AI in claims handling: The most dangerous mistake looks good
EY recently had to retract a report. One of the largest consultancy firms in the world. The Financial Times (Stephen Foley) reported that the report was full of errors produced by AI. Fabricated figures. Footnotes leading to dead pages. A reference to a McKinsey report that does not exist. Research group GPTZero found it. EY took the piece offline and is now investigating how it ever got past the checks.
The report looked fine. That was the problem.
A McKinsey report that never existed
What stands out here is not that AI makes mistakes. We know that. What stands out is that the piece looked polished and credible. Neat structure, fluent language, tidy citations. Except half of it was wrong. The figures contradicted each other, the sources did not exist, and yet it went out the door at an organisation that trades on quality.
And it is not a one-off slip. Deloitte had to amend a report for a Canadian government last year after fabricated academic sources turned out to be in it. This is happening at the firms that derive their reputation from diligence.
The problem is not the speed
Everyone talks about what AI makes faster. More output. Quicker drafts. Files summarised in a fraction of the time. All of that is true.
Almost nobody asks the harder question. If the work looks polished, structured and finished by default, does checking it not become precisely harder?
I think it does. AI raises the lower bound of the presentation. A draft that used to be shaky now looks professional. The structure is right, the tone is right, the formatting is right. But the weak assumptions are still in there. So are the misunderstood facts. They are just better dressed.
A slick assessment is not yet a correct one
Translate that to claims handling. In the past, you could tell from a file when something was off. A sloppy summary, a gap in the reasoning, an assessment that was half finished. Those were your signals. That is where you started probing.
Those signals are disappearing. A coverage assessment drafted by AI reads smoothly and assertively, even when the reasoning rests on nothing. A rejection letter confidently references a policy clause that in reality says something else. A summary quotes an amount or a date that is just slightly wrong. A fraud signal sounds plausible without anything underneath it.
It looks customer-ready. That is exactly where it goes wrong.
The mistake you only see again at Kifid
For a non-life insurer, this is not a theoretical risk. An unjustified payout costs money. An unjustified rejection costs more: a complaint, a trip to Kifid, a file you cannot explain because the reasoning rests on a fabricated source.
And the regulator is watching. You must be able to justify, trace and explain a decision. That is difficult when you no longer know exactly where an argument came from. A model that confidently invents a source does not deliver an explainable decision. It delivers a time bomb in your file.
AI with a fence around it
Here lies the difference with how we approach it at DCSolutions. What went wrong at EY happens when you let a general AI model loose on a question and allow it to generate freely. Then it can invent a policy clause or an amount, because there is nothing holding it back.
Our line is the opposite. We deploy AI in a constrained way, within fixed boundaries, on a data structure you manage and can verify. The model works only with the policy, the file and the data that actually exist. It does not invent a clause, because it only knows the clauses that are in the policy. It does not pull an amount out of thin air, because the amounts come from your own systems. Every step is traceable to a source that exists.
That constraint is exactly what makes the outcome usable in a process where you must be able to justify every decision.
Your best claims handler becomes more important, not less
There is a persistent idea that AI makes the claims handler redundant. In practice, the handler's judgement becomes more valuable, because the nature of the work shifts.
In the past you mainly caught the visible errors. Now you have to test the assumptions hidden beneath a tidy piece of work. Less checking whether it looks good. More probing whether the reasoning holds and whether the source is sound. That is exactly the craftsmanship you have your best people for, and that AI does not replace.
What you need to build now
The parties who build a serious culture of scrutiny around AI now will lead the pack later. That means constraining AI and putting it on a data structure you can verify. Whoever assumes that a slick outcome is also a correct outcome will run into themselves sooner or later. In a complaint, in an audit, or at Kifid.
Polished output is not yet good output. Certainly not in a profession where you must be able to explain every decision.
Source: Stephen Foley, ‘EY retracts study after researchers discover AI hallucinations’, Financial Times (May 2026). Freely accessible coverage at International Accounting Bulletin.