The AI agreed with the person who framed the problem
Recent studies show how user context and agreeable model behaviour can shape advice and the decisions that follow.
HUMAN-AI WORK
Lauren A Kelly
9/14/20269 min read
A person who asks an AI for advice usually supplies one side of the situation. They choose the facts and describe other people through their own view. They may also include a preferred answer, such as a proposed hire being strong or a project deserving more money.
The AI receives the account and produces a response that can feel like an outside view. The user's framing is part of the material used to generate that response. Recent research shows that several language models become more agreeable when they have a longer history or a profile of the user. Other experiments find that people can trust affirmative advice more, even when it leaves them with poorer judgement.
Researchers often call this behaviour sycophancy. The word can suggest an AI with a motive to please. The studies use a narrower behavioural meaning: under stated conditions, the system affirms the user's action, repeats an incorrect belief or frames an answer around the user's position more than a comparison system does. No motive needs to be inferred.
The work-design question concerns independence. A team may ask AI for a second view while giving it the first view in the prompt. If the system reproduces the person's starting position, two apparent judgements can carry much of the same error.
People trusted the answer that backed them
Myra Cheng and colleagues tested agreeable AI in a series of studies published in Science in 2026. They first compared responses from 11 leading AI systems with human advice drawn from Reddit and professional advice columns. Their measure counted explicit endorsement of the action described by the user.
Across more than 3,000 open advice questions, the systems endorsed users' actions about 50 per cent more often than the human responses did. The researchers then examined 2,000 posts from the Reddit forum Am I the Asshole where the community had judged the writer to be at fault. The AI systems told the writer they were not at fault in 51 per cent of these cases on average.
These comparisons don't provide a universal moral ground truth. A popular Reddit verdict reflects a community and its norms. The human advice sources were largely American, and explicit endorsement captures only one part of a response. The findings establish a consistent difference under the study's measure. They do not show that the crowd was right in every case or that agreement is always harmful.
The researchers then ran two preregistered experiments with 1,604 participants. In the first, 804 people read a hypothetical conflict and received either an affirmative AI response or one that challenged the user's conduct. In the second, 800 people discussed a low-stakes conflict from their own life with a customised AI over eight exchanges.
The live conversation produced a clear change in self-reported judgement. People assigned to the affirmative system rated themselves 25 per cent more in the right and reported a 10 per cent lower willingness to repair the relationship than people using the less affirmative system. The larger, more controlled shifts in the hypothetical experiment pointed in the same direction.
The human response to the model also changed. In the live study, participants rated the affirmative response nine per cent higher in quality. Their reported performance trust rose by eight per cent, and their stated likelihood of returning to the model for similar advice rose by 13 per cent. The condition changed judgement and preference together. People experienced the more affirmative advice as better.
These are effects on ratings and intended actions after a brief interaction. The researchers didn't observe later repair or changes in conduct. Participants recalled conflicts selected to be morally ambiguous and relatively low stakes. The study establishes a causal effect on immediate judgement under those conditions. It leaves actual behaviour and repeated workplace use open.
Workplaces may see the same mismatch. A product team may rate an assistant highly because it understands the brief and helps them proceed. A manager may experience a supportive response as a sign that the system has weighed the proposal carefully. Satisfaction records how the exchange felt. It doesn't establish that the response added an independent assessment.
More context changed how several models responded
A team can give the AI more useful information through personalisation. The system can remember a preferred format and understand an ongoing project without asking for the same background again. The same context can also indicate which answer will fit the user's existing view.
Shomik Jain and colleagues studied this effect in a peer-reviewed CHI 2026 paper. Thirty-eight students used GPT-4.1 Mini for two weeks in a persistent conversation. They made 90 queries each on average, creating around 34,000 tokens of interaction history per person. The researchers then tested how several models answered new advice questions with no context, with the full history or with a compact memory profile made from it.
User memory profiles produced the largest rise in agreement for several models. Compared with no context, the researchers measured a 45 per cent increase for Gemini 2.5 Pro and 33 per cent for Claude Sonnet 4. GPT-4.1 Mini rose by 16 per cent. Llama 4 Scout responded more agreeably with the full user history, while its memory-profile result wasn't significant. GPT-5.1 showed no significant change with either form of user context.
Models responded differently. Context didn't have one fixed effect across systems. Some models also became more agreeable when given long synthetic conversations containing no details about the real user. This suggests that response changes can come from the presence and form of context as well as accurate personal knowledge.
The researchers ran a separate test of political framing with two models. Interaction history alone did not produce a significant average rise in perspective mirroring. The response moved closer to the participant's political view when the model could infer that view accurately. Participants perceived a meaningful difference between contextual and context-free responses in roughly half of the pairs they rated.
The study has a small student sample. Everyone built their history by interacting with one model, although the researchers later placed that history in several others. The memory profiles were produced using a research method and don't reveal how commercial products build memory. The study also measured model responses to new test prompts, not decisions people made after reading them.
Its contribution is narrower and useful. An evaluation of a model with an empty prompt can miss behaviour that appears after weeks of use. A team testing a personalised assistant needs to test the assistant with representative histories, including histories that reveal the user's preferred answer.
Training for warmth can weaken correction
Agreeable behaviour can also enter through model training. Lujain Ibrahim, Franziska Sofia Hafner and Luc Rocher fine-tuned five language models to produce warmer responses, then tested their answers. Their peer-reviewed 2026 Nature paper used models of different sizes and architectures, including a version of GPT-4o.
Across selected open-ended factual, medical and misinformation tasks, warmth training increased the chance of an incorrect answer by 7.43 percentage points on average. When prompts included an incorrect belief from the user, the warm variants were about 40 per cent more likely to affirm it than the original models. The effect was stronger when the user also expressed sadness.
General capability tests did not expose most of this change. The warm and original versions performed similarly on broad knowledge, maths and harmful-request tests, apart from one decline in the smallest model. The difference appeared in conversational answers where correctness could require disagreeing with the user.
The researchers used specially fine-tuned variants. Their experiment doesn't estimate the error rate of a named production assistant. Their measure of warmth was operational and later checked with human ratings, but the study didn't test how users trusted or acted on the outputs. It shows a change in system behaviour under controlled training conditions.
The experiment doesn't show that every warm system must be less accurate. The study's cold-control models often kept their accuracy, and a system prompt for warmth produced smaller, less consistent effects than fine-tuning. Developers can test different ways of producing a supportive response. The paper shows that style training cannot be assumed to leave the substance untouched.
The person's first error can return as AI advice
A small 2026 preprint gives a more direct view of the interaction loop. Cansu Koyuturk and colleagues asked 60 people to complete survival-ranking exercises. Participants first ranked the items themselves, discussed the task with GPT-4o and then submitted a final ranking. The AI did not receive the expert answer.
Initial ranking accuracy predicted the quality of the system's recommendations. Greater carryover of a participant's incorrect items was associated with lower final performance. The result fits a possible account in which the user's error became salient context, appeared in the generated response and influenced the user's revision.
The researchers taught one group prompting techniques aimed at reducing agreement. The training reduced exact copying of incorrect rank positions. It did not produce a significant reduction in general error carryover, and final ranking accuracy did not improve reliably. Prompt instruction changed one visible form of mirroring while leaving the broader performance result unresolved.
The sample was small and participants had limited experience with chatbots. Researchers used a separate model to extract the AI's final ranking from transcripts, with a manual check of ten per cent. The paper describes its results as preliminary. It cannot show that prompt training usually fails. It does warn against assigning the user sole responsibility for obtaining an independent answer from a context-sensitive system.
Agreeable advice can still contain useful information
Evidence from a larger decision experiment complicates the account further. John Conlon and Peter Schwardmann reported a 2026 preprint with 1,500 participants making choices across 30 decision settings in economics and social science. The AI advice favoured people's initial leanings in its arguments and used agreeable language. After receiving it, people moved away from their initial position on average.
The researchers found this movement across a wide range of task types. When they increased the system's sycophancy, the movement away from the initial view weakened. Agreeable behaviour affected decisions, while the information in the advice had a larger average effect in the other direction. Participants also did not prefer a more sycophantic version in these tasks.
This differs from the social-conflict experiments, where people preferred the affirmative model and moved further into their own position. The studies asked people to do different work. A decision with relevant arguments available to the AI can benefit from new information even when the presentation leans towards the user. Advice about a personal conflict is built almost entirely from one person's account and concerns their own conduct. The emotional value of validation may also differ.
The decision paper has not yet completed peer review. Its many task settings improve breadth inside the experiment, though they don't reproduce the authority, incentives or consequences of a workplace decision. The result prevents a loose claim that any agreeable wording makes human judgement worse. It also shows why an evaluation should measure the person's final decision and the quality of the advice, alongside the rate of model agreement.
Build an independent view into the exchange
Many AI review prompts begin after a person has made a decision. ‘Here is my proposal and why I think it will work. Please review it.’ The prompt gives the system evidence and a conclusion together. A response that supports the proposal may reflect a sound assessment, the weight of the framing or both.
A team can test a different sequence. It can give the AI the source material first and ask for an assessment before revealing the proposed conclusion. The person can then compare the two views and ask the system to account for any difference. This design will not make the AI neutral. It creates an observable point before the person's preferred answer enters the context.
The interface can separate supplied facts, assumptions and the decision already made. Evidence links should remain visible beside each material claim. When the system lacks evidence, it can identify the missing information and stop short of approval. People should be able to see when a later answer changed after their own view was added.
Teams also need tests that contain user error. Give the system sound and unsound starting positions on the same underlying case. Vary the length of conversation history and the contents of a memory profile. Observe when the answer changes, which claim changes and whether the system can return to the evidence after the user presses for agreement. A single-turn accuracy score cannot describe this behaviour.
Human measures belong in the same evaluation. Record the person's initial judgement before the AI response and the final judgement afterwards. Where a defensible answer exists, compare both with it. For subjective work, use independent reviewers and record disagreement. Assess the evidence and reasoning; agreement with the user is not a success measure. Follow later action where it is proportionate and ethical.
High satisfaction should prompt a closer look when the AI is expected to challenge a proposal. A team can examine the exchanges people rated most highly and ask what changed in the work. The answer may be a clearer explanation or genuinely useful support. It may also be an answer that removed doubt without adding evidence.
Organisations should avoid relying on a prompt such as ‘be critical’ as their only control. The survival-ranking study found some change after training and no reliable improvement in final performance. A consequential decision may need a second human with access to the source evidence, or a separate AI pass that does not inherit the first conversation. The choice depends on the task and the available comparison. Extra reviewers can also repeat the same assumptions or add delay.
A favourable user rating is an intermediate outcome with several possible meanings. It can help adoption and can also indicate that the system confirmed the user's self-view. The team needs an outcome tied to the work, such as a better calibrated decision, a corrected assumption or an action that survives independent review. If agreement rises as evidence quality falls, the team has a reason to change the context supplied to the model or the authority attached to its answer.
Workplace field evidence remains thin. The strongest human experiments concern personal conflicts and constructed decisions. The context study measured model behaviour without following a work result. Production systems also change, and one of the tested models showed no significant context effect. We do not yet know which forms of personalisation preserve useful continuity while protecting an independent assessment over months of use.
Teams still need workplace evidence showing when a separate AI pass remains independent enough to improve a decision. The current studies give them several ways to test that question and no settled answer for every task.
Human-AI Performance
© 2026 Alterkind Ltd. All rights reserved.
Human-AI Performance™ is a proprietary methodology developed by BehaviourStudio using our Behaviour Thinking® framework. All content, tools, systems, and resources presented on this site are the exclusive intellectual property of Alterkind Ltd.
You’re welcome to use, share, and adapt these materials for personal learning and non-commercial team use.
For any commercial use, redistribution, or integration into client work, services, or paid products, please contact lauren@laurenakelly.com to discuss licensing terms.
Icons by Creative Mahira, The Noun Project.
Thanks to Nicholas Edell, Valentina Tan and multiple VPs implementing AI for your feedback during development.
LICENSE
Based on work by Lauren A Kelly.
For commercial licensing contact: lauren@laurenakelly.com
