What AI Actually Does in Insurance Claims
(Part 2): Where It Fails and What the Regulator Asks

Part 2 of the insurance AI walkthrough. The numbers Markel, Tractable and Daido Life publish, what continuous underwriting does to a premium, the four ways our claims stack failed, what the state examiner asked, and what to do if you are about to file a claim.

AI insurance claims processing failure modes and the questions regulators ask

Part 2 of the insurance AI walkthrough. The numbers Markel, Tractable and Daido Life publish, what continuous underwriting does to a premium, the four ways our claims stack failed, what the state examiner asked, and what to do if you are about to file a claim.

In Part 1, I walked through the claims engine we rebuilt with a mid-size general insurer: six lanes, nine model families, a fender-bender paid in seven minutes. The diagrams do not show what happens when the engine meets the real world. The real world has lawsuits, regulator inquiries, and redlining risk. It has claims that sit on the auto-approval boundary for three days because nobody wants to make the call. It has policyholders whose telematics data says one thing while their memory says another. This part covers the numbers the industry quotes, the failures we managed in the first year, and the questions the state examiner asked. It ends with practical guidance for anyone on the receiving end of an AI-evaluated claim.

Key Takeaways

  • The biggest published AI wins in insurance are workflow wins (Markel’s 113% underwriter productivity), not smarter pricing.
  • Continuous underwriting turns a once-a-year price into a stream of sensor data, and the privacy trade is only as fair as the local data law.
  • Our stack failed four ways: zip-code proxies, vague denials, bad photos, and a training set that under-represented older and imported cars.
  • The state examiner asked for process, not math: disparate-impact tests, denial reasons, vendor oversight, and who signs when the model is wrong.

What Is Inside

The Numbers That Moved

Four data points keep showing up in insurance AI decks. Before we set the insurer’s own targets I went through each one with the head of claims. Each says something different, and none says quite what the slide implies.

Free to use, share it in your presentations, blogs, or learning materials.
AI insurance numbers Markel Tractable Daido Life straight-through processing underwriting improvements
Four data points. Each is real, and each tells a different story about where AI moved the needle.

Markel with Cytora reported a 113% uplift in underwriter productivity and a cut in quote turnaround from 24 hours to 2 for strategic broker partners. Cytora is a workflow platform for underwriters, not a risk model. The gain comes from removing manual data entry and submission sorting. The lesson we took from it: the biggest AI underwriting wins are usually workflow, not smarter pricing. Our own underwriting team’s first ask was the same. Get the re-keying out of my day.

Tractable’s claim: repair estimates in minutes instead of days. Tractable is computer vision for motor damage, used by many insurers. The minutes-not-days claim is true on straight-through cases. What it does not say is that 30 to 40% of motor claims still need an adjuster. The photos are ambiguous, the damage is pre-existing, or the value is above the auto-payment threshold. Computer vision estimates fast. The hard cases stay hard. Our figure was 29% to an adjuster in year one, so right at the good end of that range.

Straight-through processing moved from 10 to 15% to 70 to 90% at mature AI insurers. That is the biggest visible customer-experience shift in decades. The 70 to 90% is motor only. Property is more like 35 to 55%, and commercial lines are still mostly adjuster-driven. The industry blended average sits around 45 to 55%, up from about 12% before AI. Our motor number was 71% and our property number 38%, and I would not trust any vendor who quotes one blended figure.

Underwriting cycle times fell from 3 days to 3 minutes on retail motor and personal lines, and fraud detection improved by around 30% on motor claims. Daido Life, a Japanese life insurer, went further and published a “glass-box” underwriting model that shows the decision logic to the underwriter beside the recommendation. That was a direct answer to regulator pressure on explainability, and it is the design we copied for our own underwriting screen.

How Our Own Year-One Scorecard Came Out

Here is the insurer’s scorecard after twelve months, in the same order the industry decks use. Motor straight-through went from 12% to 71%. Property went from 9% to 38%. Cost per motor claim fell from 140 dollars to 52. Median cycle time on a fast-track motor claim fell from 11 days to 9 minutes. The blended median across all claims, which is the number nobody puts on a slide, fell from 14 days to 4. Confirmed fraud referrals from the SIU rose 31%, and the SIU’s confirm rate on referrals rose from 22% to 41%, which is the number the head of SIU cares about most.

Two numbers went the wrong way, and both matter. Adjuster headcount stayed flat, not down, because the cases that remained took longer. And the complaint rate rose for two months after go-live, then fell below the old baseline. The two-month bump came from customers who did not trust a seven-minute payout and phoned to ask what was wrong. The fix was a sentence in the approval notice explaining that the estimate came from the photos and that a human review was one tap away. Small, and it cut the calls by half.

Continuous Underwriting and the Privacy Trade

The biggest structural change in insurance AI is the move from annual to continuous underwriting. Annual underwriting prices a policy once, at the start, on the best data available that day. Continuous underwriting keeps updating the risk view. Telematics for motor. Connected-home sensors for property. Wearables for life. If driving improves, the premium falls at renewal. If a sensor reports a leaking pipe for three weeks, the home policy may need repairs before renewal.

Here is what happens when a customer plugs in the dongle or installs the app. The device streams acceleration, braking, cornering, and GPS to the insurer. A Flink pipeline computes per-trip features: harsh-brake count, speed variance, night-driving share, urban-versus-highway mix. An LSTM trained on claim outcomes grades the resulting time series into a risk score. That score feeds a premium adjustment that appears on the next billing cycle, usually six months out. At the insurer, 23% of motor customers had opted into telematics by the end of year one, and the ones who stayed for six months saw an average premium cut of 9%.

The privacy trade is not subtle. The insurer now knows where you drive, how fast, and when. US federal privacy law does not cover telematics data the way it covers health or financial data, so consent is largely a checkbox at inception. EU customers get GDPR, with purpose limitation and deletion rights. In practice, telematics pricing is sharper where data protection is weaker, and drivers there pay less at the price of more surveillance. We wrote the insurer’s data retention policy for telematics ourselves, 24 months then delete, because nobody else in the room had an opinion and someone had to.

A fair telematics programme, from what we learned, has four traits. It is opt-in, with a plain page that shows what the device records. Customers get a copy of their own trips. Any premium increase in one renewal is capped; ours is 15%. And anyone can leave without penalty. The insurer had two of the four at launch. It has all four now, and enrolment went up after the cap was added, not down. People will trade data for a discount when they can see the deal; they will not when the deal is hidden.

What happens when the telematics data disagrees with the claim? This is the failure mode I saw most. A driver says the crash happened at 45 mph. The dongle says 68. An adjuster takes the claim and asks the driver to explain the gap, and one of three things is true. Most often the driver remembers wrongly. Sometimes the dongle is wrong because a tunnel or a garage blocked GPS, which is rare and real. And sometimes the driver is lying, which is also real. The adjuster cannot always tell which, and the model has no way to know. So the decision still goes to the human, and in our first year about 6% of telematics-linked claims needed that conversation.

The property side gave us a sharper lesson. In the second winter, the connected-home leak sensors the insurer had shipped to 12,000 policyholders started reporting moisture in a cluster of homes over one weekend. The continuous-underwriting pipeline did what we had built it to do and queued 410 repair-notice letters. A claims manager held them for a day, and it was a good thing. A firmware update from the sensor vendor had changed the humidity scale. Nothing was leaking. Every automated action that reaches a customer now waits for a human release when the volume in one day is more than three times the trailing average. That rule cost nothing to write and would have saved 410 apology calls if it had existed a week earlier.

Four Ways Our Claims Stack Failed

These are the four failures we managed in the first year. Somewhere in the industry, each one also has a regulator examination or a class-action filing behind it. That is how we knew where to look.

Free to use, share it in your presentations, blogs, or learning materials.
AI insurance claims failure modes redlining black-box denials photo quality biased training
Four failure categories. Each has been documented in a regulator examination or a class-action filing in the last three years.

1. Redlining Through Zip-Code Proxies

The underwriting model is never told about race. It sees zip code, credit score, prior claims, and vehicle make, model and year. All four correlate with race in the US to varying degrees. The model learns those correlations because they also correlate with loss experience, and the outcome is that pricing differs by race even though race was never an input. That dynamic sits behind several class actions against auto insurers in the last five years. It is also why the NAIC issued its 2023 model bulletin requiring disparate-impact testing across protected classes.

What “test” means in practice. Run the model. Compute approval rate or average premium by the racial mix of each zip code. Apply the four-fifths rule, where any group below 0.80 of the reference group is a flag. Then explain the variance or change the model. Our first run flagged three clusters of zip codes in one metro. The variance traced back to vehicle age, which is a legitimate loss factor. The pricing team still capped the weight of that feature. A legitimate reason that produces an illegitimate pattern is not a defence an examiner accepts. The test now runs monthly, and the CCO signs the output.

2. Black-Box Denials

A denial with a vague reason (“additional review required”, “inconsistent with policy terms”) leaves the policyholder nothing to appeal. Many state regulators now require specific reasons on denials, much like the adverse action notice in lending. How insurers do it varies wildly. Some insurers give a paragraph-level SHAP explanation. Some give a policy clause reference. A few give nothing. Our first denial template was in the second group, and the compliance team sent it back with the comment that a clause number is not a reason. The version that shipped names the clause, the finding, and what the customer can send to reopen the file. The NAIC bulletin and the state frameworks (Colorado SB21-169, Connecticut SB 3) are all tightening in this direction.

3. Photo Quality Failures

The pipeline assumes the photos are useful. Bad light, a bad angle, or a photo of the wrong side of the car all drop the detector’s confidence. A well-built pipeline escalates at that point. A poorly built one guesses, and often guesses wrong. Across the industry, up to 8% of motor claims at weaker insurers carry a wrong estimate because the photo-quality gate was missing or too loose. We had this failure in month two, before the overlay prompts, when a customer’s blurred bumper shot produced a “minor” grade on what turned out to be a cracked radiator support. That single claim is why the quality gate now sits in front of the damage model instead of behind it.

4. Biased Training Data

If the training set for the damage model over-represents certain makes (usually Japanese and American sedans from the last ten years), the model does worse on everything else. Older cars. Luxury brands. Fleet vehicles. Grey-market imports. The miss is silent. A confident estimate comes back 15 to 30% low for those classes. Only an accuracy audit split by vehicle class catches it, and most insurers never run one. We ran ours in month four because an adjuster complained that the model under-priced every ten-year-old pickup, and the audit agreed. We reweighted the training sample and sent under-represented classes to an adjuster until the retrain landed.

The Regulator Perimeter, and What the Examiner Asked

Three frameworks matter most for an AI insurance team in 2026, and we met one of them in person.

The NAIC Model Bulletin on the use of AI systems by insurers was issued in 2023, and 25 state departments of insurance had adopted it by early 2026. It sets governance expectations: documented model validation, adverse-action explanations, disparate-impact testing, vendor oversight. It is not a law, but state regulators use it as examination criteria.

The EU AI Act classifies insurance pricing and claims decisions as high-risk AI. EU insurers must publish transparency notices, run conformity assessments, and allow human oversight of adverse decisions, with fines up to 35 million euros or 7% of global turnover. Most EU insurers restructured their AI governance in 2024 and 2025 to meet the deadlines.

State-level US frameworks keep multiplying. Colorado SB21-169 requires insurers to test algorithms for unfair discrimination. New York DFS has issued guidance on consumer data and AI. Connecticut, California, and Washington are at various stages. The pattern is clear: state action is filling the federal silence, and an insurer operating nationally has to meet the strictest state’s bar everywhere.

When the examination came, it asked five things, and none of them was about the math. A model inventory, with an owner for each one. Disparate-impact test results for the last six months, with dates. Ten denied claims and the reason each customer received. Contract terms with the computer-vision vendor covering model updates, incident notice, and audit rights. And the name of the person who signs off when the model turns out to be wrong at scale. We had four of five. The vendor contract had no audit-rights clause, and adding one took a quarter of negotiation. A regulator does not expect zero disparate impact. They expect the insurer to have found it first.

What Changed After the Examination

Four things changed in the quarter after the examiners left, and none of them touched a model. The vendor contract gained an audit-rights clause and a 72-hour incident-notice clause. Our denial letter gained a plain-language section that names the finding and lists what a customer can send to reopen the file. Then the four-fifths report moved from a quarterly spreadsheet to a monthly job with a signature line for the CCO. And the model inventory became a living document. Each of the nine model families now has an owner, a retrain date, and a known-limits section. The engineering team grumbled that this was paperwork. The compliance team pointed out that paperwork was the entire examination. Both were right.

What I Would Tell a Chief Claims Officer

Five things, in the order I wish someone had told me. First, budget for the photo prompts and the quality gate before you budget for the vision model; the model was never our bottleneck. Second, run the vehicle-class accuracy audit in month one, not month four, because the under-estimate on old and imported cars is there from day one and only an audit finds it. Third, write the auto-payment threshold and the zero-estimate rule with the claims managers in the room. They know how a claim goes wrong. Engineers do not. Fourth, treat every automated customer contact as something that can be held: a daily volume cap and a human release button. Fifth, keep the logistic baseline for underwriting even when the gradient-boosted model beats it. The day an examiner asks how a price was set, the baseline is what you explain on the whiteboard.

None of that is exotic. The gap between a claims AI programme that survives its first examination and one that does not is almost entirely made of rules, reports, and contract clauses.

What the Adjuster Does Now

The adjuster’s job changed; it did not disappear. The diagram below maps how responsibility splits between the model, the adjuster, and the Chief Claims Officer at the insurer as it runs today.

Free to use, share it in your presentations, blogs, or learning materials.
AI insurance human roles split model adjuster chief claims officer responsibilities
What the model owns, what the adjuster owns, and what the Chief Claims Officer owns. No overlap on the hard calls.

A senior motor adjuster at the insurer now handles 25 to 45 cases a day. The easy ones, fender-benders with clear photos under the auto-payment threshold, never reach the desk. What arrives is ambiguous by definition. Boundary severity scores. Pre-existing damage disputes. Telematics that disagrees with memory. Any bodily injury, any total loss, and any claim where the fraud score rose but not far enough for SIU. One adjuster told me the job used to be 80% paperwork and 20% judgement, and the ratio has flipped.

What the adjuster does that the model cannot: phone the claimant and ask what happened. Visit the body shop to check the repair is complete. Negotiate with an at-fault party. Read medical records against one policy’s exclusions. Handle a bereaved family’s life claim with basic warmth. A model does none of these, and the insurer’s complaint rate fell in year one partly because adjusters finally had time for them.

Above the adjuster sits the Chief Claims Officer, who sets policy: auto-payment thresholds, fraud escalation rules, training requirements, retrain cadence. The CCO also carries the regulatory exposure when the model misbehaves at scale. When a state department issues a disparate-impact finding, the CCO’s name is on the response, which is why the CCO reads the monthly four-fifths report personally.

What to Do If You Are Filing a Claim

A few practical notes for the policyholder about to meet an AI-evaluated claims system, written from the side that built one.

First, follow the photo prompts. They are not arbitrary. They guide you to the shots the damage model can use. Skipping them, or taking hand-held photos of a general scene, is the fastest way to trigger low-confidence routing and a longer wait.

Second, if the estimate feels wrong, ask for adjuster review. Every responsible insurer offers it. The AI estimate is advisory for borderline cases and final only inside the auto-payment envelope. Outside that envelope you have a right to a human, so use it.

Third, if a claim is denied with a vague reason, write back and ask for specifics. Most state regulations, and GDPR in Europe, entitle you to a meaningful explanation. If the insurer refuses or sends boilerplate, complain to your state’s department of insurance. They track those complaints and use them at examination time; we saw ours in the examiner’s folder.

Fourth, understand what your telematics programme reports. Read the data-sharing agreement before you enrol. You can usually download the data the insurer holds on you, and knowing what the dongle sees is the only way to know how it will move your premium.

Fifth, photos on a claim are forever. The insurer keeps them, and the fraud pipeline compares every new claim against every old one. Perceptual hashing catches a reused photo from a claim three years ago in seconds, and it turns a small claim into an SIU file.

What Comes Next

Next in the series: firewalls. This is the domain where AI lives inside the dataplane and decisions happen in sub-millisecond windows. If banking fraud is sub-50 ms, firewall ML is sub-1 ms. The model family looks completely different too: CNNs on byte images, LSTMs on packet timing, random forests on TLS fingerprints, autoencoders for IoT baselines. Same approach: the architecture, the models, what fails, and what the security team still owns. The firewall walkthrough starts here.

Related Reading


References


Frequently Asked Questions

Does AI set my insurance premium?

AI produces a risk score. A pricing engine applies that score to a rate book approved by the state insurance department, and the final premium is the product of the two. AI does not set rates on its own; it prices risk inside a regulator-approved framework.

Can telematics raise my insurance rates?

Yes. Telematics programmes grade driving behaviour and adjust the premium at renewal. They are usually opt-in and usually offer discounts for safe driving. Harsh braking, high-speed events, and night-time urban driving raise risk scores; steady highway driving lowers them.

Can I ask for a human adjuster to review my AI claim?

Yes. Every mainstream insurer offers escalation to a human adjuster. If the AI estimate feels wrong, ask for review in writing. Most state regulations require a response within a defined window, typically 10 to 30 days depending on the state.

What is redlining in AI insurance?

Charging higher premiums or refusing cover based on geography that correlates with race or national origin. A model can redline by accident when it learns the link between zip code and risk. State and NAIC rules now require insurers to test for disparate impact across protected classes.

What is the four-fifths rule?

A test from US fair-lending law, now used in insurance: compute the approval rate or average premium for each group, divide by the rate for the reference group, and flag any group below 0.80. It is a policy check applied to model outputs, not a machine learning technique.

Why does my claim get a lower estimate than the body shop?

Often because the damage model was trained mostly on common recent sedans and under-estimates older, imported, or fleet vehicles by 15 to 30%. Ask for adjuster review and send the shop’s itemised estimate; well-run insurers route under-represented vehicle classes to a human until the model is retrained.

Does the EU AI Act apply to insurance?

Yes. It classifies insurance pricing and claims decisions as high-risk AI. EU insurers must publish transparency notices, run conformity assessments, and allow human oversight of adverse decisions. Fines can reach 35 million euros or 7% of global turnover.