ELU NOVA
← All field notes
The pilot-to-production gap

The Klarna reversal and the cost of unwinding

4 min read

The most-cited AI deployment of 2024, and its sequel

In February 2024, Klarna launched an AI customer service assistant built in partnership with OpenAI. The numbers the company published were striking and were repeated everywhere: within its first month the assistant handled 2.3 million customer conversations, work Klarna equated to roughly 700 full-time agents, with approximately $40 million in annualised savings reported in its Q1 2024 materials.

For roughly a year it was the reference case for AI replacing operational headcount. Klarna went from around 5,000 employees in late 2023 to approximately 3,500 by late 2024, largely through attrition under a hiring freeze — though it is worth noting that not all of that reduction is fairly attributed to the assistant, since the company was also under general cost discipline following the 2022–23 fintech downturn.

In May 2025, CEO Sebastian Siemiatkowski told Bloomberg the company had gone too far. Klarna reopened hiring for human agents. His stated reasoning was direct: the company had focused too heavily on cost, and quality suffered as a result.

The reversal is frequently written up as a story about the limits of AI. We read it differently, and we think the more useful reading is less comfortable.

The volume was measured. The quality was not.

Klarna measured throughput with precision. 2.3 million chats. 700 agent-equivalents. Two-thirds of conversations automated. Those numbers were tracked closely and reported publicly within weeks.

What does not appear to have been running at equivalent rigour, from the start, was customer-outcome measurement segmented by case type.

This matters more than it sounds. In conversational AI deployments, aggregate satisfaction can hold steady while satisfaction among a small, high-value cohort of complex cases collapses. The average conceals the failure. Disputes, complex refunds and financial hardship situations are exactly the interactions where the cost of a poor outcome is highest and the volume is lowest — which means they contribute little to the aggregate and a great deal to the brand.

Klarna's assistant absorbed the routine tier well. On the evidence of the CEO's own account, the problem was that the value tier had been absorbed too, and the measurement in place did not surface that quickly enough.

Separate research points at the same mechanism from a different angle. MIT's Project NANDA, in The GenAI Divide published in July 2025, attributed widespread pilot stalling to systems that do not retain feedback, adapt to context or improve over time — a "learning gap". Klarna had no learning gap on volume. It had one on outcome.

Unwinding is more expensive than not doing it

The part of this case that belongs in a business case, and almost never appears in one, is the cost of reversal.

Klarna's correction required recruiting, onboarding and training customer service staff to rebuild capacity that had been allowed to decay through attrition. It also meant absorbing a quality gap that had already reached customers before it was identified — a cost that does not appear on any line item and does not disappear when the staff are rehired.

There is a further cost that is genuinely difficult to recover. Experienced agents carry implicit knowledge: how to navigate a particular edge case, what a customer in a specific situation usually actually needs, when a routine-looking query is a symptom of something larger. That knowledge left with the people. Rebuilding headcount is a hiring problem. Rebuilding that knowledge takes years.

AI replacement business cases typically model labour cost savings. They rarely model the revenue impact of a satisfaction decline, the churn cost from customers who left over service experience, or the cost of unwinding the strategy if it underperforms. The Klarna case suggests those omitted variables can exceed the modelled savings.

What Klarna did not do, and what it did

Two things are worth being precise about, because the story is often told carelessly.

Klarna did not abandon AI. The assistant remained in production on the high-volume routine tier. What changed was the addition of a human path for the cases where the AI had not held parity. The company described the new model as hybrid, with remote agents on flexible schedules, working alongside AI tooling rather than being replaced by it.

And Klarna's workload was, in AI terms, close to a best case: high-volume, structured, authenticated consumer-fintech intents, with direct engineering access to a frontier model provider. Organisations working with ambiguous intents and unstructured inputs — B2B, insurance claims, healthcare, complex logistics — should expect a lower ceiling on automation and a higher permanent human-in-the-loop floor, not a similar one.

What to take into your own deployment

Measure outcome and volume from day one, at the same rigour, segmented by case type. An aggregate satisfaction number will not show you a collapsing high-value cohort until the damage is done. Segment before you scale, not after.

Keep humans on exceptions from the start. Not as a transitional phase to be removed once confidence rises, but as a permanent design element with a defined routing rule. Klarna's hybrid model is where it ended up; it would have been cheaper as a starting position.

Do not let capacity decay ahead of proof. A hiring freeze is a bet that the automation will hold. If it does not, the reversal costs more than the freeze saved, and the institutional knowledge does not come back with the headcount.

Model the unwind cost. Ask what it would cost to reverse this if it underperformed, and put the number in the business case. If nobody can produce it, the case is incomplete.

Klarna's error was not deploying AI. It was deploying it instead of people rather than alongside them, and measuring cost more carefully than quality. Both are avoidable, and both are cheaper to avoid than to correct.

Find out where AI actually pays off in your business.

A focused discovery audit with a senior principal. You leave with a prioritized, costed roadmap — whether or not you work with us.

contact@elunovatech.com · Response within one business day