Designing the handover after an AI failure
The latest field evidence from customer service shows how the history of an AI-led exchange and a worker's available attention shape the handover.
HUMAN-AI WORK
Lauren A Kelly
8/1/20267 min read
Customer-service handovers carry their own history. By the time a person enters an AI-led chat, the customer may have repeated the problem, waited through several unsuccessful turns and become frustrated. The worker receives the original request along with the effects of the service so far.
This makes the familiar phrase 'human in the loop' too loose for work design. It places a person somewhere in the process without specifying when they enter, what information reaches them, how many cases they are watching or how much authority they have once they take over.
A May 2026 working paper by Yiwei Wang and colleagues gives us an unusually detailed view of this problem. The researchers studied an agentic AI system on Alibaba's Taobao platform. It spoke directly with customers and tried to settle a limited set of standard service requests. Human workers watched the AI-led chats, could step in on their own judgement and also received alerts when monitoring software detected a technical problem or customer frustration.
The study follows work in a live service operation. Its evidence concerns one organisation, one short deployment and a narrow set of customer requests. The latest version is a preprint, so it hasn't completed peer review.
Late handovers were harder to recover
The field experiment involved 647 remote gig workers and 680,676 chats. Workers were randomly assigned to continue with all-human service or to supervise the agentic AI while handling other chats themselves. The treatment ran for 17 days in August 2024. Only 5.8 per cent of the chats in the analysis were classed as suitable for the AI.
Across all chats, the AI treatment cut average chat duration by 3.2 per cent. It produced no statistically significant overall change in customer ratings or in the rate at which customers returned with the same issue within seven days. A manager looking only at the headline figures could reasonably report a small speed gain and stable quality.
Grouping the results by chat eligibility changes the interpretation. AI-eligible chats became 16.8 per cent faster, while their customer ratings fell by 0.412 points on a five-point scale. The human-only chats handled by workers in the AI treatment became 1.8 per cent faster and received slightly higher ratings. The researchers link this spillover to workers switching away from those cases less often and giving them more attention.
AI changed the distribution of attention. Some customers received a fast automated service that they rated less favourably. Other customers benefited because a worker had more time for them. The unchanged average rating contains both experiences.
The researchers also examined 11,069 AI-led chats and matched each one with a similar all-human chat. A technical alert usually meant that the request had moved beyond the AI's capability. After a person took over, those chats lasted 19.1 per cent longer than comparable all-human cases, while ratings and repeat contact were statistically similar. Service quality was preserved in these matched cases, at the cost of extra time.
An emotional alert came after the monitoring software detected frustration or doubt. Those chats lasted 40.8 per cent longer than comparable all-human cases. Repeat contact rose by six percentage points and ratings fell by 0.928 points. Workers also replied more slowly after these escalations. They sent fewer messages and did less information seeking and solution work. Their measured empathy was similar to the other escalation groups.
People who stepped in on their own judgement tended to act earlier. Their cases still lasted 9.5 per cent longer than matched all-human chats and ratings fell by 0.524 points. The fall was smaller than it was after an emotional alert. Repeat contact was 3.2 percentage points lower, although this estimate had weaker statistical support.
The authors argue that earlier intervention preserved the worker's ability to put the service back on course. The data fit that account. The study design doesn't establish the timing claim as causal. Workers were randomly assigned access to the AI, while the type and timing of escalation were observed after assignment. The researchers matched cases on twelve measured features. Unmeasured differences could still explain part of the pattern.
Customer ratings also came from 17 per cent of chats. Response rates stayed similar across the experimental groups, which helps, although the people who chose to rate may differ from those who didn't. The workers were paid by volume, handled an average of 3.14 chats at once and combined AI supervision with direct customer work. Each feature could affect when they noticed trouble and how much effort they could give it.
Random assignment gives strong field evidence for the overall effect of this deployment. The matched subgroup analysis gives suggestive evidence about the effect of handover timing. That distinction matters when we move into design.
The person receives a changed task
Monitoring an AI-led conversation is a different job from handling the customer request at the start. The worker must notice weak signs across several live chats, judge when the AI has reached its limit and take responsibility for an exchange they didn't begin. An alert may arrive after the customer has already formed a view of the service.
The AI also changes the material that the person works with. It has selected questions and offered explanations. These actions shape the customer's replies. A technically correct summary cannot remove this history. The person may need to repair trust before the customer will provide another detail or accept a proposed solution.
Research on assistive AI shows why the AI's position in the service matters. In a peer-reviewed field experiment with a meal-delivery company, Shunyuan Zhang and Das Narayandas gave some human agents access to AI-generated suggestions. The assistance generally improved response speed and customer sentiment, especially for less-experienced staff. The AI worked behind the person, who continued to speak with the customer.
The same study found a problem after an earlier chatbot had failed to understand the customer. If the transferred human then used AI assistance, very fast replies could lead customers to believe they were still dealing with the bot. Customer sentiment worsened. The handover changed the meaning of the next response. Speed, which helped elsewhere, became a cue that the failed interaction was continuing.
A second 2026 Alibaba field experiment studied an assistant that offered diagnoses and possible solutions near the start of a chat. Average service speed and customer ratings improved. Lower-performing workers gained most. Top performers received worse customer ratings and saw more customers return with the same problem. The process data suggest that using the tool interrupted how these workers managed concurrent chats. This paper is also a preprint, and its explanation of the mechanism is suggestive.
These studies examine different arrangements. One AI acts before the person and speaks to the customer. Another sits behind the person and offers material they can use or ignore. Both arrangements redistribute attention and alter the sequence of work. Their effects also vary with the worker's skill and the customer's recent experience.
A 2024 meta-analysis in Nature Human Behaviour covered 106 experiments and 370 effect sizes. Human-AI combinations performed better than people working alone on average. They performed worse than the better of the human-only or AI-only condition on average. Results varied greatly across studies, and the review covered research published only up to June 2023. The presence of both contributors therefore tells us little about the performance of a current arrangement.
Design the handover as part of the work
A team can begin with the point at which the AI should stop. Capability limits belong in that decision, along with evidence that the exchange itself is deteriorating. Repeated requests for the same information, several unsuccessful solution attempts and a change in customer sentiment may each justify intervention. Human workers also need a usable way to take over before the monitoring software reaches its threshold.
The threshold creates a workload for someone. Earlier alerts may make cases easier to recover while drawing workers away from other customers. Alert volume, supervisory capacity and response time need to be designed together. A person assigned to watch too many conversations can remain formally responsible while having little practical chance to intervene.
Ruth Schmidt and Zeya Chen's conceptual work on positive friction in human-AI interaction is useful here. A prompt that asks a worker to review a case adds effort and interrupts their flow. That friction may serve a purpose when it creates a timely moment for judgement. Their conceptual paper supports a testable proposition: compare the cost of an interruption with any improvement in recovery.
The transfer itself needs to carry enough state for action. The worker should be able to see the customer's original request, what the AI tried, any commitments it made and the reason for escalation. They also need to know which facts remain uncertain. A generated summary can help, provided the worker can inspect the relevant exchange and supporting records. The interface should mark the moment a person joins so the customer understands who is now responding.
Workers also need enough authority after the transfer. A person who inherits a damaged exchange may need permission to correct an AI claim, offer a remedy or change the normal service route. If the role allows only an apology and another scripted answer, an earlier alert will do little.
Teams need to measure the work beyond speed. In the Alibaba trial, the overall average concealed lower ratings for AI-eligible chats and better ratings elsewhere. Teams should examine resolution and customer experience for each important case type, alongside worker attention. They should also look for changes by worker experience. Longer-term observation can test whether supervision builds useful judgement or leaves people with fewer chances to practise the underlying task.
For an AI service, a team might revise an escalation threshold when late emotional alerts exceed an agreed rate, when workers cannot respond within the required time or when a case type shows a sustained fall in resolution. Changes to the model, customer demand and staff experience can each alter a threshold that once worked. Recording the reason for each revision also creates operational knowledge that a single launch measure will miss.
Evidence we still need
The Alibaba study ran for 17 treatment days. It cannot tell us how supervision changes skill, fatigue or judgement over several months. It gives limited evidence about services where mistakes carry legal or clinical consequences. The eligible chats were selected because they involved standard requests, so the results should not be carried across to open-ended casework without local testing.
Researchers could now randomise the handover rule itself. One group might receive earlier alerts based on lack of progress, while another uses the existing emotional threshold. Experiments could vary the number of chats each person supervises and the information shown at transfer. Researchers should report AI and worker behaviour, the sequence between them and the final result.
Teams can test smaller changes in live work now. They can describe the current service, record its results by case type and state exactly how AI will alter the allocation of work. They can then observe what people and AI actually do, compare the result with the baseline and revise the design. This establishes whether adding AI made that piece of work better for the customer and for the person responsible for finishing it.
Human-AI Performance
© 2026 Alterkind Ltd. All rights reserved.
Human-AI Performance™ is a proprietary methodology developed by BehaviourStudio using our Behaviour Thinking® framework. All content, tools, systems, and resources presented on this site are the exclusive intellectual property of Alterkind Ltd.
You’re welcome to use, share, and adapt these materials for personal learning and non-commercial team use.
For any commercial use, redistribution, or integration into client work, services, or paid products, please contact lauren@laurenakelly.com to discuss licensing terms.
Icons by Creative Mahira, The Noun Project.
Thanks to Nicholas Edell, Valentina Tan and multiple VPs implementing AI for your feedback during development.
LICENSE
Based on work by Lauren A Kelly.
For commercial licensing contact: lauren@laurenakelly.com
