Skip to content
Verinika
Fairness by Design: Testing AI in Insurance Underwriting and Claims
Back to Insights
Best Practices

Fairness by Design: Testing AI in Insurance Underwriting and Claims

AI is transforming how insurers assess risk, price policies, and process claims. But when algorithms determine who gets coverage and at what price, the potential for discrimination is enormous. This deep analysis examines the regulatory framework, the tension between actuarial accuracy and fairness, and the assurance practices that can make insurance AI both effective and equitable.

February 19, 2026
Verinika Team
21 min read

The insurance industry has always been in the business of discriminating—in the actuarial sense. Distinguishing between risk profiles is the foundation of pricing and underwriting. But there is a crucial difference between actuarially justified differentiation and unlawful discrimination, and artificial intelligence is making that line harder to see and easier to cross.

AI systems now assess applications for life and health insurance, calculate premiums, detect fraudulent claims, and automate coverage decisions. When these systems are well-designed and properly validated, they can improve pricing accuracy, accelerate claims processing, and reduce human error. When they are not, they can systematically disadvantage protected groups, deny coverage to people who need it most, and expose insurers to regulatory sanctions that can reach tens of millions of euros.

The Regulatory Landscape

The EU AI Act: Insurance as High-Risk

The EU AI Act (Regulation 2024/1689) classifies AI systems used for risk assessment and pricing in life and health insurance as high-risk under Annex III. This is one of the clearest regulatory signals that insurance AI is subject to enhanced scrutiny. The high-risk designation triggers a comprehensive set of obligations that apply from August 2, 2026.

As a high-risk deployer, an insurer must implement a quality management system (QMS) to track model design, training data, and performance benchmarks. It must ensure that training, validation, and testing datasets are representative and relevant. It must provide meaningful human oversight, with trained staff who can identify anomalies and override automated decisions. And it must offer clear explanations to customers when AI influences decisions regarding policy acceptance, coverage, or pricing.

The penalties for non-compliance are severe. Violations of prohibited AI practices can attract fines of up to 35 million euros or 7% of global annual turnover. Non-compliance with high-risk obligations can result in fines of up to 15 million euros or 3% of turnover. These are not theoretical maximums—they are designed to be genuinely deterrent.

Prohibited Practices: The Social Scoring Ban

The EU AI Act also contains outright prohibitions that may constrain certain data practices in insurance. Most notably, social scoring—evaluating individuals based on data collected in unrelated contexts to produce detrimental treatment—is prohibited. This could restrict insurers from using certain external consumer data sources if the data originates from contexts unrelated to insurance risk.

For example, an insurer that factors social media activity or online shopping behaviour into risk assessment may find that such data falls within the scope of the social scoring prohibition if the resulting treatment is disproportionate or the data context is deemed unrelated. The boundaries here are still being tested, but the direction of travel is clear.

Interaction with Existing Frameworks

The EU AI Act does not operate in isolation. Insurers must continue to comply with Solvency II (governance, risk management, capital adequacy), the Insurance Distribution Directive (IDD—fairness, transparency, acting in the customer's best interest), DORA (ICT risk management and operational resilience), and the GDPR (data protection and rights related to automated decision-making, including the right to meaningful explanation under Article 22).

The interaction between these frameworks creates a complex compliance landscape. An insurer deploying AI must satisfy the AI Act's technical requirements while also meeting the IDD's conduct-of-business standards and the GDPR's data subject rights. This requires coordinated governance rather than siloed compliance efforts.

Fairness by Design: Testing AI in Insurance Underwriting and Claims

The Fairness-Accuracy Trade-Off

When Fairness Costs Capital

One of the most important—and least discussed—tensions in insurance AI is the relationship between fairness constraints and capital requirements. Research has shown that applying strict fairness constraints such as demographic parity to actuarial models can increase capital requirements by up to 4.1% of own funds.

This happens because strict demographic parity can force a model to deviate from pure risk-based pricing, creating cross-subsidisation between risk groups. From a Solvency II perspective, this deviation introduces additional uncertainty that must be covered by capital buffers. The result is that fairness is not free—it has a measurable financial cost.

This does not mean fairness should be sacrificed to capital efficiency. It means that the trade-off must be acknowledged, quantified, and managed. Insurers need to understand how different fairness definitions affect their pricing, reserving, and capital models, and make informed decisions about which trade-offs are acceptable.

Beyond Demographic Parity

Demographic parity—ensuring that outcomes are proportionally equal across groups—is only one definition of fairness, and it is often not the most appropriate for insurance. Equalised odds (equal true positive and false positive rates across groups) may be more relevant for claims fraud detection. Calibration (equal accuracy of predictions across groups) may be more appropriate for risk assessment.

The choice of fairness metric should be driven by the specific context and the potential harms. A claims fraud system that has a higher false positive rate for one demographic group imposes a disproportionate burden of investigation on that group. A risk assessment model that systematically over-predicts risk for one group charges them more than their actual risk warrants. Different harms call for different fairness metrics.

Common Failure Modes

Proxy Variables and Indirect Discrimination

Insurance AI is particularly susceptible to proxy discrimination. A model that does not directly use protected characteristics can still discriminate through correlated variables. Geographic data can proxy for ethnicity. Occupation can proxy for gender. Digital footprint data—browsing patterns, device usage, app installations—can proxy for income, education, and demographic characteristics in ways that are difficult to detect without deliberate analysis.

The challenge is compounded by the sheer volume of features that modern machine learning models can process. A traditional underwriting model might use a dozen variables. A modern ML model might use hundreds. As the number of features increases, so does the risk that subtle proxies for protected characteristics are present in the model without anyone being aware.

Data Quality and Representativeness

Insurance models are trained on historical data that reflects historical underwriting and claims patterns. If certain populations were historically under-insured—because they were less likely to apply, more likely to be rejected, or priced out of the market—then the training data underrepresents their risk profiles. A model trained on this data may perform poorly for these populations, either by mispricing their risk or by failing to detect legitimate claims patterns.

Claims Automation Bias

On the claims side, AI systems that automate or triage claims processing introduce risks of systematic error. If a claims model is trained on historical decisions that reflected adjuster bias—for example, more sceptical treatment of claims from certain geographic areas or demographic groups—the model will learn and perpetuate that bias at scale.

Building an Insurance AI Assurance Programme

Pre-Deployment: Bias Testing and Documentation

Before deploying any AI system in underwriting or claims, insurers should conduct comprehensive bias testing across all relevant protected characteristics. This means measuring outcomes—acceptance rates, pricing levels, claims approval rates, fraud referral rates—across demographic groups, not just in aggregate.

The results must be documented as part of the technical documentation required by the EU AI Act. This documentation serves a dual purpose: it provides evidence of compliance for regulators and creates a baseline against which ongoing performance can be measured.

Human Oversight: Meaningful, Not Mechanical

The EU AI Act requires human oversight for high-risk AI systems. But "human oversight" is meaningless if the humans involved lack the training, information, and authority to meaningfully challenge the AI's outputs. Effective human oversight means providing underwriters and claims handlers with AI explanations they can understand, clear criteria for when to override the AI, training on the specific risks and failure modes of the system, and protection from pressure to simply accept the AI's recommendation for efficiency reasons.

Post-Deployment: Continuous Monitoring

Insurance portfolios are not static. Customer demographics change, economic conditions shift, fraud patterns evolve, and regulatory expectations increase. A model that was fair and accurate at deployment can drift into unfairness or inaccuracy over time.

Continuous monitoring should track underwriting and claims outcomes across demographic groups, compare AI decisions to human-only benchmarks, monitor for feature drift and population stability, and alert when predefined fairness thresholds are breached.

Vendor Due Diligence

Many insurers use AI models developed by third-party vendors—particularly for fraud detection, telematics analysis, and digital underwriting. The EU AI Act and supervisory expectations make clear that using a vendor's model does not transfer risk or responsibility. Insurers must conduct due diligence on vendor models that is at least as rigorous as the validation they would apply to internally developed models.

The Bottom Line

AI in insurance is not inherently unfair—but it is not inherently fair either. Fairness requires deliberate design, rigorous testing, continuous monitoring, and genuine human oversight. The EU AI Act's classification of insurance AI as high-risk, combined with existing requirements under Solvency II, IDD, and GDPR, creates a comprehensive but complex compliance landscape.

Insurers that treat fairness as a design principle—not a compliance afterthought—will build AI systems that are both commercially effective and genuinely equitable. Those that treat it as a box to tick will eventually face the consequences, both regulatory and reputational.

Ready to Evaluate the Systems You Deliver?

Tell us what your team is building and what must be demonstrated before the next client or release decision.

Discuss a Partner Pilot