How Disparate Impact Shapes AI Pricing in Insurance

How proxy variables and outcome differences appear in insurance AI pricing, what testing can establish, and where state-specific procedures belong.

In this article

For Compliance officers, actuaries, model-risk teams, and counsel interpreting fairness results in insurance pricing and underwriting.

Read if You need the concepts behind proxy and outcome testing before applying a particular state's procedure.

By Simon Li · Published JUL 14, 2026 · Updated SEP 25, 2026 · 6 min read

Engraved cover illustration: How Disparate Impact Shapes AI Pricing in Insurance
Ask AI

The governance map separates general control design from the legal and analytical concepts used in particular decisions. In pricing, a model can be blind to race, sex, disability, or another protected characteristic and still produce uneven results. The model may use geography, credit, education, occupation, household information, or behavior that correlates with group status. Many weak signals can also combine into a stronger proxy no analyst intended to create.

Examining that problem takes a handful of concepts: disparate treatment, disparate impact, proxy variables, outcome differences, business necessity, and less discriminatory alternatives. It does not reproduce New York’s numbered procedure or Colorado’s filing regime. Those belong to the New York NYDFS Circular Letter 7 and Colorado guides.

Treatment, impact, and proxy are different ideas

Disparate treatment concerns intentional different treatment because of a protected characteristic. Disparate impact describes a neutral practice that has a disproportionate adverse effect on a protected group. The familiar burden-shifting structure comes from federal employment law: a challenger shows the effect, the decision-maker offers business necessity, and a less discriminatory alternative may still matter.1

Insurance statutes and regulatory documents do not always use the phrase “disparate impact.” They may refer to unfair discrimination, disproportionate adverse effects, bias, or protected-class correlation. The vocabulary and legal consequence depend on the jurisdiction and line. The analytical structure remains useful because it separates three questions:

  1. Was protected-class status used directly or was someone treated differently because of it?
  2. Does a neutral input or derived feature correlate with protected-class status?
  3. Does the complete model produce materially different outcomes across groups after relevant risk differences are considered?

A proxy is a mechanism, not a legal conclusion. A variable can correlate with a protected group and still have a legitimate insurance relationship. A model can also show outcome differences even when no single input has a strong correlation. That is why input and output analysis should not be collapsed.

How proxy effects enter a pricing model

Pricing models work by finding patterns. The same capability that improves prediction can reconstruct social and economic differences from ordinary-looking data.

Pricing-data families, their proxy risks, and the questions to examine
Data familyWhy it can matterWhat to examine
GeographyHistorical segregation, infrastructure, hazard, service access, and local cost can overlapDirect risk relationship, granularity, neighborhood error, alternatives
Credit and financial dataPayment behavior and financial stress may correlate with loss and with protected groupsPermitted use, stability, causal story, incremental lift, group effect
Occupation and educationWork and educational opportunity are unevenly distributedInsurance relevance, categorical treatment, missingness, alternatives
Telematics and behaviorCollection and opportunity to change behavior varyExposure normalization, device bias, access, discount and surcharge effects
Property imageryImage quality and housing stock differ by area and property typeError by cohort, inspection correction, filed treatment, nonrenewal effect
Third-party scoresA single score can hide many sources and transformationsSource disclosure, version, validation population, limitation, audit evidence

Exhibit D of the NAIC AI Systems Evaluation Tool helps identify categories of data and their sources.2 It does not ask whether a data element is a proxy. Exhibit C asks about testing outputs for matters that include accuracy, unfair trade practices, and unfair discrimination.2 The Evaluation Tool guide walks through the form. The fairness team brings the proxy and outcome questions to the data and model record.

Input testing and outcome testing answer different questions

An input screen asks how strongly a feature, source, or derived score tracks protected-class status. It can find an obvious proxy and help prioritize review. It cannot show what the whole model does after interactions, transformations, and offsets.

An outcome analysis compares decisions or prices across groups, with careful attention to legitimate risk and product differences. It can reveal a model-level disparity. It cannot by itself identify the responsible feature or settle whether the difference is legally prohibited.

Both analyses depend on design choices:

  • which population and time period are included;
  • how protected-class information is obtained or estimated;
  • which outcome is measured, such as quote, declination, tier, premium, renewal, or surcharge;
  • what constitutes a similarly situated comparison;
  • how small samples and missing data are handled;
  • which thresholds trigger investigation rather than an automatic verdict.

The method, limitation, and decision made from the result belong in the record. A single metric without those elements is easy to overread.

Business necessity is a documented argument

When a feature or model produces concern, the insurer should be able to explain what business need it serves and why that need is legitimate, lawful, and fair under the governing regime. Predictive lift is relevant, but “the model performs better” is incomplete. Performance may be concentrated in one segment, depend on unstable data, or add little value after a less problematic variable is used.

A less discriminatory alternative asks whether the business objective can be met with another feature, transformation, threshold, model, or workflow that reduces the disparity. Useful comparisons preserve the business target and report the tradeoff. An alternative that destroys the intended risk signal is not equivalent; an alternative rejected because it is unfamiliar has not been seriously tested.

The analysis should preserve candidates considered, evaluation method, performance and group results, the decision, and the person authorized to make it. New York specifies where and how this search fits its process; the Circular Letter guide carries those details.

Three regulatory contexts, kept separate

Four regulatory instruments and the limits of what each contributes
InstrumentWhat it contributesWhat it does not do
NAIC Model BulletinProgram expectations for detecting and addressing unfair discrimination and managing AI riskPrescribe one statistical test or create a new statute3
NAIC Evaluation ToolData-source and selected-model questions a regulator can useUse the term proxy or define the insurer’s high-risk criteria2
New York Circular Letter 7State-specific expectations for ECDIS, AIS, outcome assessment, alternatives, governance, transparency, and noticeCreate a generic procedure for every state or every insurance function4
Colorado SB 21-169 and Regulation 10-1-1Colorado insurance authority, governance framework, records, and reporting for the covered scopeCopy New York’s language or establish one nationwide metric5

This comparison prevents a common documentation error: writing a blended “regulator requirement” that no regulator actually issued. Where each state’s own text departs from the model bulletin is set out on the New York and Colorado state pages.

A review that can be defended

Start with one pricing or underwriting model and create a compact review file:

  1. Identify the decision, population, product, jurisdictions, model version, and material data sources.
  2. Screen likely proxy features and record the method and limitations.
  3. Compare relevant outcomes across groups, controlling for legitimate differences with actuarial and legal review.
  4. Investigate the features, interactions, thresholds, or workflow steps that may explain a material difference.
  5. Evaluate business necessity and credible alternatives under the applicable standard.
  6. Record the decision, conditions, owner, monitoring handoff, and state-specific filing or notice work.

The model monitoring playbook covers production thresholds and drift response. The vendor playbook covers third-party diligence. Keeping those controls separate prevents this fairness review from becoming another copy of the entire governance program.

What a result can and cannot prove

A correlation is a signal to investigate. It is not a verdict. A clean aggregate result may hide a subgroup or interaction problem. A statistically significant result may be operationally small, while a large result may be too imprecise to interpret. Estimated protected-class data introduces its own uncertainty.

The defensible conclusion states the limits plainly: what was tested, what was not, which legal and actuarial judgments were applied, and what happens next. Whether a particular model violates law is a fact- and jurisdiction-specific legal question. Testing supplies evidence for that judgment; it does not replace it.

Footnotes

  1. Civil Rights Act of 1964, Title VII, as amended by the Civil Rights Act of 1991, 42 U.S.C. § 2000e-2(k). ↩

  2. National Association of Insurance Commissioners, AI Systems Evaluation Tool 4.0, 2026. ↩ ↩2 ↩3

  3. National Association of Insurance Commissioners, Model Bulletin on the Use of Artificial Intelligence Systems by Insurers, adopted December 4, 2023. ↩

  4. New York State Department of Financial Services, Insurance Circular Letter No. 7 (2024), July 11, 2024. ↩

  5. Colorado Division of Insurance, SB21-169 insurance AI resources and Regulation 10-1-1. ↩

The Bottom Line

  • A protected-class field can be absent while a model still reconstructs group differences through correlated variables.
  • Input correlation and outcome difference are separate tests. Neither result, standing alone, proves that a model is lawful or unlawful.
  • Business necessity and the search for a less discriminatory alternative are part of the reasoning, not a license to skip measurement.
  • New York and Colorado use different legal instruments and scopes; their detailed procedures should not be blended into a national checklist.

Recommended next

Who Owns the Evidence in Insurance AI Governance

Insurance AI governance roles as an ownership matrix: which function prepares each piece of NAIC evidence, which one signs it, and who answers for it in an exam.

Continue →
Engraved portrait of Simon Li

Written by

Simon Li · Founding Editor

I write InsureAI Wire and maintain its 51-jurisdiction tracker. Most of the work is reading: NAIC working group papers, state bulletins, bills, court filings, and public comment letters. Every claim on the site carries the document it came from, so you never have to take my word for it.

Contact or report a correction →

Related reading

Information aggregation and analysis, not legal advice. See our disclaimer.