← BlogBlog

Enterprise AI agent deployment case studies - what six known rollouts actually show

Search for an enterprise AI agent deployment case study and you get press releases with the walk-back missing. The launch post arrives in month one with a big number on it. The correction arrives a year later in an interview, and the two are never on the same page. Below are six deployments read through the same five questions, each with a source you can open: what was deployed, what the company claimed, what is actually confirmed, what went wrong afterwards, and what transfers to your own rollout. One of the six is ours, treated the same way as the rest.

6Deployments dissected
2.3MKlarna chats, month one
245MFargo interactions, 2024
58%Best agent, single-step CRM

Read every enterprise AI agent deployment claim through the same five questions

Most deployment stories are published by the company that ran the deployment. That is not dishonesty, it is timing: the launch post gets written in month one, and by the time anyone can tell whether the thing held, the news cycle has moved on. The useful habit is to separate the claim from the confirmation and to go looking for what happened next.

Five questions do most of the work. What was actually deployed, and to how much of the business? What did the company claim, in its own words? What is confirmed by someone other than the company? What went wrong after the launch post? And what part of it applies to a business that is not theirs? The fourth question is the one press releases skip, and the fifth is the only one that changes what you do on Monday.

It also helps to know what a competent agent scores before your process gets involved. Salesforce's own research team published CRMArena-Pro, a benchmark of nineteen expert-validated CRM tasks, and the frontier models of the day cleared around 58% of single-turn tasks and about 35% once the task ran across several turns. The same paper reports near-zero inherent confidentiality awareness in the agents, and notes that prompting them into it costs task performance. Those numbers are not a verdict on any product, but they are a useful floor: when a vendor says its agent handles 80% of tickets, the interesting question is what counts as handling one.

What a good agent scores before your process is involvedSource: CRMArena-Pro, Salesforce AI Research, arXiv:2505.18878 (retrieved Aug 10, 2026)
Workflow execution, single turnthe narrowest task type
over 83%
All nineteen tasks, single turn
~58%
All nineteen tasks, multi-turnthe shape most real support work takes
~35%

One more habit before the cases. A deployment that got rolled back is not proof the technology does not work, and a deployment still running is not proof it does. Read the direction of travel, not the headline. We wrote separately about the failure modes that survive a passing demo, and most of what follows is those failure modes with company names attached.

Klarna: the number was real, the conclusion was not

What was deployed. An OpenAI-powered customer service assistant, live across Klarna's markets and fronting the main support channel rather than sitting in a pilot corner. This is the deployment everyone in the industry has argued about, which is exactly why it is worth reading from the source.

What was claimed. Klarna's own press release from February 2024 is specific: 2.3 million conversations in the first month, two-thirds of all customer service chats, available in more than 35 languages across 23 markets, errands resolved in under 2 minutes against 11 minutes previously, a 25% drop in repeat inquiries, and an estimated $40 million profit improvement for 2024. It also says the assistant was doing "the equivalent work of 700 full-time agents".

What is confirmed. The volume and the timing, by Klarna. Everything else is Klarna measuring Klarna. Note the phrasing on that famous figure too: the release says equivalent work of 700 agents, not that 700 people were replaced. The paraphrase that spread across the internet is not what the company wrote.

What went wrong. By May 2025 the CEO was walking it back in public. Klarna reopened hiring for human customer service, with Sebastian Siemiatkowski saying cost had been "a too predominant evaluation factor" and that the result was lower quality, alongside a commitment that a customer who wants a human can always reach one. Quality on the hard tail of cases was the thing nobody had a number for.

What transfers. The deployment was not the mistake, the objective function was. Cost per contact and quality of resolution move in opposite directions on complex cases, and if your only measurement comes from the agent's own logs, you cannot tell the two apart. Pick a resolution metric you can verify without the bot's cooperation before you scale anything. We laid out which support metrics survive contact with reality in what AI customer support ROI actually looks like.

Air Canada: your chatbot is not a separate legal entity

What was deployed. A customer-facing support chatbot on the airline's website. Modest by 2026 standards, which is the point: the legal principle it produced applies to anything you put in front of a customer, agentic or not.

What happened. A passenger booking travel after a death in the family was told by the bot that he could apply for the reduced bereavement rate retroactively, within 90 days of the ticket being issued. Air Canada's actual policy said the opposite. He paid full fare, applied later, and was refused.

What is confirmed. All of it, by a published decision. In Moffatt v. Air Canada, 2024 BCCRT 149, the British Columbia Civil Resolution Tribunal found negligent misrepresentation and awarded the difference between the fare paid and the bereavement fare. Air Canada's defence was that it could not be held liable for information given by its own chatbot. The tribunal rejected that flatly: a chatbot has an interactive component, but it is still just part of the website, and it makes no difference whether the information comes from a static page or a bot.

What transfers. Whatever your agent says is a statement your company published. That reframes the design problem: the answers that create an obligation are a small, listable set, and they should not be free-form generation at all. Prices, refunds, eligibility, delivery dates and warranty terms come out of retrieved approved text or they do not come out. Everything else can be conversational. In the EU there is a second layer on top of the liability question, since disclosure duties around AI systems interacting with people are now live - we mapped the current position in the EU AI Act compliance calendar.

Drive-thru voice: one chain quit, the other narrowed

What was deployed. Automated voice order taking at the drive-thru, by two of the largest restaurant groups in the world, in the noisiest and most latency-sensitive customer environment anyone has tried this in.

What was claimed, and what is confirmed. McDonald's ran automated order taking with IBM from 2021 and ended the test in 2024, switching the technology off in participating restaurants by 26 July that year. The company's public line was that the test gave it confidence voice ordering would be part of its restaurants' future, and it did not publish accuracy figures, cost figures, or the metrics it used to judge the pilot. So the honest summary is: the partnership ended, the goal did not, and the viral videos of mis-taken orders are anecdote rather than data.

What went wrong at the other chain. Yum went in the other direction and hit the same wall from a different angle. After putting voice AI into around 500 restaurants in early 2025, it slowed the rollout that August following customer complaints and a wave of deliberate trolling, including an order for 18,000 cups of water. The public conclusion from Yum's technology side was not that voice AI fails, but that busy restaurants may still be better served by a person taking the order.

What transfers. Two things, and neither needs a restaurant. First, adversarial input is a load case, not an edge case. The moment a system is public, some share of its traffic exists to break it, and a quantity field with no sanity bound is the cheapest possible failure. Second, "where do we not deploy this" is a legitimate answer that deployments rarely plan for. Site-by-site or queue-by-queue scoping beats an all-or-nothing rollout, and it gives you a control group for free.

Commonwealth Bank removed 45 roles, then put them back

What was deployed. A generative AI voice bot in the contact centre of Australia's largest bank, alongside a decision to declare 45 customer service roles redundant on the strength of the call volume it was expected to absorb.

What went wrong. The finance sector union challenged the redundancies and the underlying claim. According to the account published by the Australian Computer Society, incoming call volumes were actually rising, to the point where the bank was offering overtime and having team leaders answer calls. In August 2025 CBA reversed the decision and apologised, stating that its initial assessment that the 45 roles were not required "did not adequately consider all relevant business considerations and this error meant the roles were not redundant".

What is confirmed. The reversal and the bank's own wording, which is unusually direct for a public company. What was never established is the deflection figure the original decision rested on.

What transfers. Deflection is a measurement, not a projection, and it has to be taken on total queue volume over a comparable window rather than on the share of sessions the bot closed by itself. A bot that ends 30% of chats while total contacts rise 20% has deflected nothing. The people best placed to notice that are the ones answering the overflow, which is an argument for asking them before the headcount decision rather than after.

The deployment that scaled kept the model away from the data

What was deployed. Fargo, Wells Fargo's assistant inside the mobile banking app, handling everyday tasks: balances, transfers, transaction lookups, bill payments. The scope is deliberately narrow and the volume is very large, which is probably why it got a fraction of the coverage Klarna did.

What was claimed. Growth from 21.3 million interactions in 2023 to 245.4 million in 2024, and roughly 336 million cumulative as of reporting in April 2025. Bank figures, so the same caveat applies as everywhere else on this page.

What is interesting is not the number. It is the data path. Speech is transcribed locally, the text is scrubbed and tokenized inside the bank's own stack, a small model detects personally identifiable information, and only sanitized intent and entities are sent out to the language model. Detokenization happens back inside. The bank's CIO Chintan Mehta described the arrangement as being the filters in front of and behind the model.

What transfers. In that design the language model is an intent classifier sitting in front of a deterministic execution layer, and the thing that actually moves money is ordinary software with ordinary controls. That is the reliability answer and the compliance answer at the same time, which is why it keeps showing up in regulated deployments. It is also the practical version of the distinction between an agent and a workflow: most of what people call an agent is a fixed path with a model doing the language part, and it costs less and breaks less precisely because of that. The wiring question, how the model reaches live systems without holding a copy of them, is its own piece.

Our entry: a voice agent with a two-second budget

What was deployed. A voice race engineer for live iRacing sessions. The driver speaks over push-to-talk, Whisper transcribes, an OpenAI agent reads live telemetry through 19 MCP tools across six categories, and the answer comes back through TTS in the style of a pit radio call.

What was measured. Answers in under 2 seconds, telemetry for 64 cars, a worker polling the simulator every 100 ms and writing completed laps to a local database the agent queries on demand. These are our numbers from our own build, which puts them in the same self-reported bucket as everyone else's on this page.

What went wrong. The obvious architecture, keeping the full tool-call chain in conversation history the way a chatbot keeps a transcript, poisoned the next answer. A lap time captured 30 seconds ago is already wrong, and the model cannot distinguish a stale number from a fresh one because both are just text in the context. The fix was to stop treating history as storage: a four-message sliding window holds only the question and the final answer, every intermediate tool call runs fresh each turn and is discarded, and live state lives outside the model entirely.

What does not transfer. It is a simulator, one driver at a time, no personal data, no regulatory surface. It says nothing useful about a bank's controls or a contact centre's staffing. What does transfer is the pair of constraints: a hard latency budget forces you to decide what gets retrieved at question time rather than remembered, and high-risk actions get a gate. Reading telemetry is autonomous; pit commands pass through a human confirm before anything reaches the car. Every deployment above that got into trouble had an action path with no equivalent gate.

Where each deployment stands as of August 2026

This is the section that dates, and it dates fast. Everything below was checked against the linked sources on 10 August 2026. Open them before quoting any of it. The whole point of this piece is that the follow-up outlives the announcement.

Status check, 10 August 2026
DeploymentStartedWhere it standsConfirmed by
Klarna AI assistantFeb 2024Dual track. AI on routine volume, human support rebuilt for complex and premium cases after the May 2025 correctionKlarna press release; CX Dive (May 2025); CX Today (Jul 2026)
Air Canada support chatbotPre-2023Decided against the airline; the ruling stands as the reference point for chatbot liability*Moffatt v. Air Canada*, 2024 BCCRT 149
McDonald's automated order taking (IBM)2021IBM partnership ended, technology switched off by 26 Jul 2024; voice ordering still stated as a goalRestaurant Dive (Jun 2024)
Yum drive-thru voice AI2023Slowed Aug 2025 after complaints and trolling, then expanded with a new vendor to 890+ US restaurants across 38 statesRestaurant Dive (Jul 2026)
CommBank contact centre voice botLate 202445 redundancies reversed and apologised for in Aug 2025; the bot itself was not withdrawnInformation Age / ACS (Aug 2025)
Wells Fargo Fargo2022Running at scale; 245.4M interactions reported for 2024, PII tokenized before the modelAIM Media House (Apr 2025)

Two of the six had a public reversal. None of the six shut the technology down and went back to how things were before, which is the detail that usually gets lost when a rollback is reported. In one case the vendor changed, in another the scope narrowed restaurant by restaurant, and at Klarna the humans came back on the harder tier of cases. The software stayed in production through all of it.

What actually transfers to your own rollout

Nothing above is a template you can copy. What is portable is the set of questions each one answers badly, and you can answer them for your own deployment in an afternoon with no vendor in the room. If a question below has no owner and no number, that is the part of your rollout most likely to end up in someone else's article.

Seven questions, and the deployment that shows why each one matters
QuestionHow to answer itWhere it showed up
What does "handled" mean in your number?Pull 50 random closed sessions and read them. Count the ones where the customer's problem is actually gone, not the ones where the bot sent the last message.Klarna
Which answers create an obligation?List them: price, refund, eligibility, dates, warranty. Serve those from retrieved approved text, never from free generation.Air Canada
What happens under hostile input?Spend an hour attacking your own system: nonsense, abuse, absurd quantities, an accent it has not heard. Bound every numeric field.Yum drive-thru
Where does the deflection number come from?Total queue volume, same window before and after, taken from the queue rather than the vendor dashboard.Commonwealth Bank
What does the model actually see?Draw the data path end to end. For every piece of customer data that reaches the model, ask whether it has to.Wells Fargo
What is the latency budget, and what is stale?Set the number the answer must beat, then mark which inputs are retrieved fresh per turn instead of carried in context.Race Engineer
Which actions are gated?List the irreversible ones - payments, cancellations, anything sent to a customer - and put a human confirm in front of each.All six

One concrete next step, if you want a single one: take last month's real conversation log, pull fifty at random, and mark each as resolved, escalated or abandoned by reading it rather than by trusting the status field. That number is your actual baseline, and in our experience it rarely matches the dashboard. Everything else in a deployment plan is easier to argue about once it exists.

We build these systems for a living, so the bias in this piece is worth stating plainly: we think most of the trouble above was scoping, not modelling. If you want a second read on where a specific deployment is likely to break, a short audit gets you the two or three failure modes most likely to bite, mapped to the questions in the table. The table works fine without us either way.

FAQ

Is there an enterprise AI agent deployment case study with independently audited numbers?

Very few, and none of the famous ones. Almost every figure in circulation comes from the company that ran the deployment, measured on its own logs, published at the moment it wanted the story told. The exception in this roundup is Air Canada, where a tribunal decision put the facts on the public record because someone disputed them. Treat that asymmetry as the default and read announcements accordingly.

Did Klarna really replace 700 people with AI?

Not according to its own press release, which says the assistant was doing the equivalent work of 700 full-time agents. That is a workload comparison, not a headcount action, and the difference matters because the widely repeated version implies a layoff the release never describes. What is documented is the later correction: in May 2025 the company reopened hiring for human support and said cost had been too dominant a factor in the original design.

What is the most common reason these deployments get pulled back?

Capacity removed ahead of proof. Klarna and Commonwealth Bank both acted on a projected reduction rather than a measured one, and both had to reverse part of it. The model was not the failing component in either case. The failing component was a number that had not been verified on total volume before it was used to make a staffing decision.

Does a rollback mean the technology was the wrong choice?

Usually not, judging by what these companies did next. McDonald's ended a vendor relationship and kept voice ordering as a goal. Yum slowed down, changed supplier and came back at a larger footprint with a view on which restaurants suit it. Klarna kept AI on routine volume and rebuilt human capacity for the hard tier. The visible pattern is scope correction, not abandonment.

Changelog
  • 10 August 2026Published.
Free process audit

See what this would look like in your operations.

Get in touch

30 minutes · we map your 3 best automation opportunities · no obligation