AI works differently across the same job

Field studies in customer service show that one AI assistant can produce very different results across a team and across the working day. This article examines what that variation means for the way organisations divide work.

HUMAN-AI WORK

Lauren A Kelly

7/2/20266 min read

An organisation buys an AI assistant for its customer service team. A few months later, the figures look encouraging. Replies are faster and customer ratings have edged up.

Those averages describe the whole team. They conceal the variation inside it.

A request to cancel a subscription gives the AI a familiar problem with a fairly clear route through it. A third complaint about missing grocery deliveries is a different kind of case. The customer has a history, the failure may sit elsewhere in the company and another polished apology may make things worse. The cases arrive in the same queue and are handled by someone with the same job title. They place very different demands on the person and the AI.

This makes the work-design question more specific: at what level should an organisation divide work between people and AI? A decision made at role level will often be too coarse. The useful dividing line may sit between case types, between stages of a case or between people at different points in their experience.

What the research found

Several large field studies now show how much variation can sit inside a positive average.

In Generative AI at Work, published in the Quarterly Journal of Economics in 2025, researchers studied the introduction of an AI assistant among 5,172 customer support agents. Access to the assistant increased the number of issues resolved per hour by 15 per cent on average. Less experienced and lower-skilled agents improved both their speed and the quality of their work. The most experienced and highest-skilled agents made small gains in speed and saw small declines in quality.

The benefits were also larger for relatively rare problems. This finding is easy to miss. People often assume that AI will be most useful on the common, repetitive cases. In this setting, it gave newer agents access to patterns that experienced colleagues had acquired through time on the job. Rare cases were exactly where that borrowed experience had more value.

A separate randomised field experiment, published in Management Science, studied online customer service at a meal-delivery company. AI-assisted agents replied faster, engaged more deeply and improved customer sentiment. Again, less experienced agents gained most.

The result changed with the conversation. AI helped with subscription cancellations. It was least effective when customers returned with repeat complaints caused by wider service failures. The assistant could improve the wording of a reply. The operation that had produced the complaint sat beyond its reach.

The study also captured something that a productivity dashboard would struggle to show. Customers first dealt with an automated chatbot. Some were then transferred to a human agent. When the chatbot had already failed to understand them, very fast replies from an AI-assisted human reduced customer sentiment. Customers appeared to infer that they were still dealing with the bot.

Speed carried a social meaning. The same behaviour that looked efficient in one conversation looked inattentive in another.

A 2026 field experiment at Alibaba gives us a third view. It is a large, randomised study of after-sales support. It remains a preprint and hasn’t yet been through peer review. The AI drafted an initial diagnosis and a proposed response. Agents could use it, change it or ignore it.

On average, the assistant reduced the time needed to identify the issue and shortened chats. Customer ratings improved. The rate at which customers came back with the same problem, used as a more objective quality measure, didn’t change significantly.

Lower-performing agents gained most. Among the top-performing agents, the researchers found little improvement in speed and declines in both customer ratings and repeat-contact outcomes. Their analysis suggested a possible reason. Experienced agents using the AI spent longer shifting their attention between concurrent chats. Replies slowed and more customers abandoned the conversation or returned soon afterwards. These measures are consistent with workflow disruption. They don’t prove the full causal chain.

The pattern across the studies is fairly clear. AI performance can vary by the worker’s experience, the customer’s history, the type of case and the measure used to judge success. An average across all four can hide the place where service has deteriorated.

The wider evidence supports this caution. A 2024 meta-analysis brought together 106 experiments and 370 effect sizes. Human-AI combinations performed better than people working alone on average. They performed worse than the better of the person or the AI working alone. The pooled difference against that stronger benchmark was small, with a Hedges’ g of -0.23. Many organisations still assume collaboration will produce the best result by default.

The studies in that review varied a great deal, and most were completed before the current generation of workplace assistants. They can’t tell a contact centre how to route tomorrow’s cases. They do show why the routing question needs to be asked.

The behaviour inside the result

The AI in these studies behaves like a store of available patterns. It can recognise a familiar shape and produce a plausible response at speed. That can give a new agent a useful starting point. It can also bring consistency to parts of the job where people would otherwise spend time searching through guidance or asking a colleague.

Its limits also have a shape. The assistant sees the information made available in the conversation and in its supporting data. A repeat complaint may depend on a warehouse problem or a promise made on a previous call. Fluent language can cover the absence of that wider understanding. The reply may sound complete before the case is complete.

People bring a different set of behaviours. A newer agent is still building a mental library of cases. Suggestions can shorten that process and give them language they haven’t yet learnt to produce under pressure. An experienced agent may already have an efficient way to scan, diagnose and move between conversations. Adding a suggestion creates another object to read and assess. A tool designed as assistance can interrupt a routine that was already working well.

Customers are part of the combined performance too. They arrive with expectations formed by everything that has happened earlier in the service. A rapid answer after a good handover can feel responsive. A rapid answer after a bot has repeatedly misunderstood the problem can confirm the customer’s suspicion that nobody is listening.

This is why the phrase “human in the loop” tells us so little. It says that a person is present. It doesn’t tell us what that person knows, what the AI has seen, how the work is moving around them or how the customer reads the interaction.

How I would design the work

I would begin by breaking the role into the smallest units that show a meaningful difference. In a service team, that may include the customer’s intent, the number of previous contacts, the source of the failure and the agent’s experience with that type of case. The aim is to find where AI assistance changes the result. The team doesn’t need a perfect map of the whole job before it can begin.

The first version will be incomplete. It should be treated as a working allocation that changes as the team learns. A cancellation request may begin in an AI-supported route. A second contact about an unresolved operational failure may move into a route where the agent sees the customer history before any suggested wording appears. New case types should remain visible until there is enough evidence to place them.

Measurement also needs to stay close to the work. “Productivity” is too broad on its own. A team can close more chats and create more return contacts. Higher customer ratings may accompany an unresolved underlying problem. The Alibaba study found movement in subjective service quality without a significant change in its objective repeat-contact measure. Both belong in the account.

For a service operation, I would watch handling time, resolution, repeat contact and abandonment. I would then split those measures by case type and agent experience. Averages remain useful after the differences have been examined.

The most capable staff need particular attention. Many rollouts use them as a reference group, a source of good prompts or informal trainers for everyone else. The research suggests that the tool can impose a cost on these workers even when it helps the team overall. Their existing routines should be observed before the interface changes. If the assistant adds reading, switching or checking, that extra work needs to appear in the evaluation.

Small points of friction can help at the cases where context is likely to matter. After a failed chatbot exchange, for example, the human agent could receive a short handover and a clear warning before an AI response is shown. A repeat complaint could require the agent to confirm the cause and the promised action before sending polished text. This adds effort in selected places. Applying it to every reply would probably create another blanket process and slow down straightforward work.

These are design proposals drawn from the pattern in the evidence. The studies don’t establish a universal routing method, and they don’t show that a particular warning or pause will improve every service. A team would need to test those changes against its own cases.

The boundary should also move. Agents learn. Models change. A case that was rare becomes common after a product fault. A new policy can turn a simple request into one that needs discretion. Review points are part of the operating process, not a final meeting held after the rollout has been declared successful.

Many customer-service teams already use AI. They now need to know where a particular assistant helps a particular person with a case, and what happens to the customer after the reply has been sent. Current research gives us good reasons to expect different answers across the same working day.

Research sources

Human-AI Performance

By Lauren A Kelly

© 2026 Alterkind Ltd. All rights reserved.
Human-AI Performance™ is a proprietary methodology developed by Lauren A Kelly using my Behaviour Thinking® framework. All content, tools, systems, and resources presented on this site are the exclusive intellectual property of Alterkind Ltd.

You’re welcome to use, share, and adapt these materials for personal learning and non-commercial team use.

For any commercial use, redistribution, or integration into client work, services, or paid products, please contact lauren@laurenakelly.com to discuss licensing terms.

Icons by Creative Mahira, The Noun Project.

Thanks to Nicholas Edell, Valentina Tan and multiple VPs implementing AI for your feedback during development.

LICENSE
Based on work by Lauren A Kelly.

For commercial licensing contact: lauren@laurenakelly.com