The AI may be training the person too

Repeated AI advice can alter the judgement of the person using it, including when they reject some of its suggestions. This article examines feedback loops between people and models, and what organisations need to measure after a system goes live.

HUMAN-AI WORK

Lauren A Kelly

4/2/20266 min read

Most organisations assess an AI assistant one output at a time. They record the accuracy of a summary, the adviser’s response to a recommendation, the number of cases completed and the time taken.

The person using the system is usually treated as a fixed part of the test. Their judgement is the reference point against which the AI is checked. Repeated interaction can change that reference point.

A housing repairs officer who sees an AI risk score on every case may gradually adjust their sense of what counts as urgent. A recruiter repeatedly shown the same pattern of recommended candidates may come to see that pattern as ordinary. Neither person has to follow every suggestion for the system to influence how later cases feel.

Evaluating the assistant one output at a time cannot capture this. The work also needs to reveal whether people are learning useful regularities or absorbing the system’s errors as they use it.

The reviewer does not stay still

A study published in Nature Human Behaviour in 2025 examined this process across a series of experiments involving 1,401 participants. The researchers studied perceptual, emotional and social judgements. People made repeated estimates and interacted with an AI system or with other people.

Small biases in human data were amplified by the AI. Participants who then interacted with the biased system became more biased in their own judgements over time. The amplification was greater than it was in human-to-human interaction. Participants also underestimated how much the AI had influenced them.

The direction depended on system quality. Interaction with an accurate AI improved people’s accuracy, and the improvement grew over repeated trials.

The researchers tested the effect across several kinds of judgement and included a widely used text-to-image system. This breadth gives the result weight. The experiments still took place under controlled conditions and used tasks that are much narrower than a job. They don’t tell us how far a claims handler or clinician will shift after a year with a workplace model.

The proposed explanation has two parts. Machine-learning systems can take a small pattern in their data and reproduce it in a stronger form. Their answers also tend to be less variable than human answers. A repeated, consistent judgement can look like a stable norm. People absorb some of that norm, then bring the changed judgement to the next interaction.

People already learn regularities through exposure. Someone who reviews hundreds of applications develops a feel for the usual case and the detail that deserves a second look. AI advice joins that stream of experience. Its consistency can give it more influence than one colleague’s opinion, even when nobody has formally granted it more authority.

People are also poor observers of their own adjustment. Participants in the feedback-loop experiments underestimated the AI’s influence. A worker can reject some recommendations and still have their threshold changed by the pattern across all the others.

Disagreeing with the AI has a cost

The 2025 paper measured this loop directly. A 2026 experiment in Harvard Data Science Review shows how the design of review work can make people more likely to leave AI errors in place.

In that study, 2,784 participants checked figures that an AI had extracted from company greenhouse-gas reports. They could accept a suggested value or flag it as wrong. The researchers changed the quality of the first few suggestions, the effort needed to correct an error and the financial reward for accuracy.

When flagging an error also required participants to type the correct value, people made fewer corrections and accepted more incorrect suggestions. The additional work sat on the act of disagreeing with the AI. A bonus for high accuracy didn’t meaningfully improve performance. The accuracy of the first three suggestions also had no meaningful effect.

People’s existing attitudes were more predictive. Participants who were sceptical of AI detected errors more reliably and achieved higher accuracy. Those with more favourable views of automation accepted incorrect suggestions more often. Scepticism won’t always improve performance. An accurate system can help, as the feedback-loop study showed. The reviewer and the review process both contribute to the observed error rate.

The extra step in this experiment was small. It still changed whose judgement survived. A pause can give someone time to form an independent view. A correction form can place the burden on the person who disagrees. Both create friction, with very different effects on the work.

Carryover needs a longer measure

Field studies give us signs that interaction can carry forward in useful ways. In a study of 5,172 customer support agents, access to an AI assistant increased productivity by 15 per cent on average. The researchers also found evidence that agents learnt from the tool and improved their English fluency, especially among international workers. Less experienced staff benefited most.

A separate randomised meal-delivery study found that AI-supported agents engaged more deeply with customers and improved customer sentiment. Less experienced agents again gained most. The study measured service performance, so it can’t tell us what people retained after using the assistance.

The 2026 Alibaba preprint found another kind of carryover within a task. Its assistant offered a diagnosis and draft response only at the opening stage of a service chat. Treated agents continued to respond faster and take a more proactive role later in the conversation, after the direct assistance had ended. Customers supplied less input. The experiment shows that an early AI contribution can alter the human exchange that follows. It doesn’t establish lasting learning.

The wider meta-analysis of 106 human-AI experiments helps place these findings. More than 95 per cent of the systems in its dataset had a person make the final decision after seeing AI input. The combined result still fell below the better of the person or the AI alone on average. Human sign-off didn’t guarantee that the team had retained the strongest judgement.

The AI has no stable behaviour outside its data, model and instructions. It can amplify a pattern because its learning process turns regularities into predictions. A model update can change those regularities overnight. The person may then be learning from a moving source without being told that it has moved.

The combined behaviour develops across a sequence. The AI offers a judgement. The person accepts, edits or rejects it. That response affects the work and may also shape the person’s next judgement. In systems that learn from local decisions, the response can later return as training data. A loop can exist even when the deployed model stays fixed, because the human part of the arrangement is changing.

What a six-month evaluation needs to see

I would evaluate the sequence as well as the individual decision. An initial trial can record accuracy, speed, overrides and the types of error made. A longer evaluation can ask whether people’s unaided judgements change after several weeks of use. Without that second measure, useful learning and harmful drift look the same in a dashboard: both happen inside the person.

One practical method is to keep a small set of cases where staff record an initial judgement before seeing AI advice. Sampling keeps this step away from most tasks. The organisation can then see whether confidence and error thresholds are moving. Known-answer cases can help with calibration, provided they resemble the work closely enough to be meaningful.

The review interface should make challenge easy. A worker who spots an error needs a quick way to flag it, correct it when they know the answer and route it when they don’t. Requiring a full reconstruction of the answer every time may reduce correction, as the 2026 experiment found. The effort of investigation can sit with the team responsible for the system once a credible concern has been raised.

Deliberate pauses belong where automatic acceptance carries a serious cost. A high-risk recommendation might remain hidden until the person has recorded a short initial assessment. Another design could surface cases where the person and model disagree, then ask for evidence before either view is adopted. These are proposals to test. The cited studies didn’t compare these exact workplace arrangements.

Override data needs a direction. A single override rate mixes useful correction with avoidable human error. Teams should examine cases where staff changed a correct AI answer, cases where they rescued an error and cases where both missed the problem. Changes in those patterns over time are more informative than a target that simply rewards agreement.

Model monitoring and staff development also belong in the same review. If an update changes the distribution of recommendations, people need to know. If experienced reviewers begin converging on the model’s distinctive errors, retraining the model alone won’t immediately restore their earlier judgement. Calibration sessions may be needed, using examples selected independently of the system.

The same process can capture beneficial learning. The customer-support study suggests that AI can spread useful language and case knowledge to newer staff. An organisation can test whether people retain that learning on cases completed without assistance. It can also check whether the tool is improving judgement or helping people produce a good answer they can’t later reproduce.

Review points should be tied to evidence. A change in model version, a new case type, a revised policy or a shift in human error patterns gives the team a reason to revisit the division of work. The response might be a different interface, more independent practice or narrower use of AI on that part of the task. Keeping every control forever would make the work cumbersome and encourage people to bypass it.

Current research gives us strong evidence that repeated AI interaction can change human judgement in controlled tasks. We also have field evidence of learning and behavioural carryover. Long-term workplace evidence remains thin. That uncertainty is a reason to measure people over time, because the human performance observed on launch day may no longer be the human performance present six months later.

Research sources

Human-AI Performance

By Lauren A Kelly

© 2026 Alterkind Ltd. All rights reserved.
Human-AI Performance™ is a proprietary methodology developed by Lauren A Kelly using my Behaviour Thinking® framework. All content, tools, systems, and resources presented on this site are the exclusive intellectual property of Alterkind Ltd.

You’re welcome to use, share, and adapt these materials for personal learning and non-commercial team use.

For any commercial use, redistribution, or integration into client work, services, or paid products, please contact lauren@laurenakelly.com to discuss licensing terms.

Icons by Creative Mahira, The Noun Project.

Thanks to Nicholas Edell, Valentina Tan and multiple VPs implementing AI for your feedback during development.

LICENSE
Based on work by Lauren A Kelly.

For commercial licensing contact: lauren@laurenakelly.com