The AI workflow became another task
A field experiment at Gap shows what can happen when an organisation prescribes exactly how people should work together with AI.
HUMAN-AI WORK
Lauren A Kelly
8/25/202610 min read
Once an organisation gives staff access to generative AI, guidance tends to follow. A product team writes prompt templates. A transformation team sets out a recommended sequence for using the tool with colleagues. The intention is sensible. Unstructured use can produce weak briefs and inconsistent checking.
Every added instruction also becomes part of the job. It takes time and depends on the supporting software. People have to remember a new sequence. When the method involves a colleague, their calendars and ability to connect enter the task too. An evaluation of human-AI performance needs to count this work.
A 2026 field experiment by Alex Farach and colleagues gives us a useful, awkward case. The researchers worked with 388 Gap employees during a one-day AI training event. All participants had access to Microsoft Copilot. The experiment changed the way pairs were asked to use it.
Each pair had 30 minutes to produce a one-page AI adoption plan for its part of the business. Ninety-seven pairs could work as they normally would, including using Copilot in any way they chose. Another 97 pairs were given a method called Create-Out-Loud. They had to meet on Microsoft Teams, discuss the plan aloud, create a transcript and ask Copilot to write the first draft from that transcript.
The prescribed method was associated with fewer finished documents and lower scores. The study has serious limitations. Its evidence concerns the performance of this particular arrangement under the conditions in which people were asked to carry it out.
The documents that never arrived
Ninety-three of the 97 pairs using their normal methods submitted a document. Seventy-one of the 97 pairs assigned to Create-Out-Loud did so. Twenty-six pairs in the prescribed group produced no document, compared with four pairs in the other group.
I keep returning to those missing documents. For a timed work task, completion is part of performance. A workflow that produces no output for more than a quarter of the assigned pairs has created an operational result, even if researchers later call it attrition.
The submitted documents also differed. The control group averaged 15.63 points on a 22-point rubric. The prescribed group averaged 10.68. About 69 per cent of the control pairs reached the study's 70 per cent quality threshold. Around 34 per cent of the prescribed group reached it.
The document history adds some texture. Control pairs produced 3.28 versions on average, compared with 2.44 in the prescribed group. This pattern fits an account in which the meeting and transcript used up time that could have gone into revising the draft. The study didn't test that explanation directly.
Many pairs couldn't carry out the new method as intended. Among the 62 treatment pairs with usable pair-level data about their behaviour, 23 couldn't meet synchronously. The researchers called them ‘Stranded’. Seventeen pairs met and then used Copilot separately, a pattern labelled ‘Parallel Play’. Twenty-two pairs reported meeting and using Copilot jointly.
These categories were made after the experiment using self-reports. They cannot establish what each way of working did to quality. They do show the routes people took when a formal process met the practical conditions of a 30-minute session. Some couldn't get the joint task started. Others kept the meeting and dropped the shared use of AI.
The word ‘compliance’ appears throughout the paper. For work design, we need to know whether people could carry the method out. The protocol bundled a live meeting with transcription and AI drafting. Trouble at the meeting stage could stop the later stages. People who changed the route may have been responding to the clock, the technology or the needs of the task. The published analysis doesn't let us separate those reasons.
The protocol changed the exchange
Create-Out-Loud altered the prompt and the work around it. It changed how two people formed a shared view of the task and passed their knowledge to the AI. The method also shaped when they first saw something they could edit.
The people had to turn their knowledge into a spoken conversation before Copilot received it. The transcript became the brief. This route may surface areas of agreement and disagreement between colleagues. It may also leave out knowledge that a person would add when writing directly, such as the name of an internal system or a constraint that feels too obvious to say aloud. The task rubric explicitly rewarded this kind of specificity. The paper doesn't analyse the transcripts in enough detail to tell us which effect occurred.
Copilot then took the role of first drafter. The method gave pairs less freedom to decide where the AI could help. A pair using the tool naturally could ask for possible risks, write its own outline or use AI to improve one difficult section. A fixed transcript-to-draft route placed the model at the centre of composition. The average number of versions suggests that the prescribed pairs did less revision, though it doesn't show how much human checking took place inside each version.
The published analysis concentrates on document production. It gives little account of the AI's behaviour. We don't see a detailed analysis of the prompts or the first drafts. We also don't know how the model used the colleagues' information. This leaves several possible mechanisms open. Copilot may have produced generic plans because the transcripts lacked business detail. People may have accepted an early draft because the clock was running down. They may also have spent time repairing a draft that followed the conversation closely but didn't fit the required document. Further study would need to test these explanations; this experiment didn't establish them.
This distinction matters when an organisation tries to apply the result. The human behaviour is partly visible: pairs failed to connect, worked separately or completed the joint route. The AI behaviour remains largely inside a black box in the reported analysis. The performance result came from the whole exchange, including the meeting software, transcription and time limit.
What the experiment can support
The size of the difference deserves attention. Its cause is harder to isolate.
Every control pair was assigned to the morning session. Every prescribed pair was assigned to the afternoon. Any difference between those sessions is therefore bundled with the change in method. Time of day may have played a part. Technical conditions or the way staff ran the event may also have differed. The authors ran sensitivity analyses and compared behaviour within the prescribed group. Those checks cannot separate the protocol from all the other session conditions.
The study wasn't preregistered. The authors clearly label several later analyses as post-hoc, including the behaviour categories and checks prompted by missing documents. This transparency helps a reader judge the evidence. It also means those analyses should be treated as investigation of an unexpected result.
The quality measure needs similar care. GPT-4o-mini graded each document three times and the researchers used the median score. Human raters checked a sample and preserved the direction of the result, although agreement between the LLM and people was imperfect. The AI grader also gave higher marks to longer documents. Score and word count had a strong correlation of 0.646. Control documents averaged 740 words and prescribed documents averaged 454. Once the researchers adjusted for word count, the estimated quality gap fell by a third and remained negative.
The authors used a statistical method that tests how selective missing outputs could alter the quality result. The estimated gap stayed negative under those assumptions. The morning and afternoon confound remains. Statistical adjustment cannot recover a comparison that the experiment never made.
All five authors work for Microsoft, and the tool in the study was Microsoft Copilot. The paper discloses this affiliation. It also describes employee consent and reports that no institutional review board oversaw the work. The affiliation and ethics statement belong in an appraisal of a company-linked working paper that hasn't completed peer review.
The paper tested another form of guidance later in the day. People in the treatment session were taught to treat AI as a thought partner, and the control session received standard training on Copilot. The prespecified comparison of continuous quality scores wasn't statistically significant. The AI grader gave 68 per cent of all documents the maximum score, which left little room to detect a difference. A post-hoc test of perfect scores favoured the treatment, although its result could be explained by the different rates of missing output. I haven't used this second task as evidence that one kind of guidance works better.
The defensible claim is quite specific. A synchronous, transcript-led AI protocol was associated with worse completion and lower document scores during this Gap training event. The experiment doesn't isolate the protocol as the sole cause. It offers no general verdict on structured AI guidance in other settings.
A working method has to fit the work
A second 2026 field experiment produced a different pattern. In the peer-reviewed Cybernetic Teammate study, 791 Procter & Gamble professionals worked on real product-development problems during a one-day virtual workshop. Researchers randomly assigned them to work alone or in pairs, with or without generative AI.
Individuals using AI produced work of a similar quality to pairs working without it. People using AI also developed proposals that crossed the usual boundary between commercial and technical expertise. Teams with AI were more likely than solo workers without AI to produce an idea in the top tenth of all submissions.
The researchers gave every AI group a set of suggested prompts. Only 38 per cent used them. Participants who didn't use the prompts performed at a similar level to those who did, and both groups did better than the conditions without AI. People chose whether to use the suggestions, so the causal effect of prompt guidance remains unknown. The observed gains were present among participants who followed the supplied technique and among those who didn't.
Both studies took place in short employee workshops using business tasks. Their tasks and working methods differed. P&G participants had a full day and worked on long-standing product questions from their own units. Gap pairs had 30 minutes to create an adoption plan and the prescribed group depended on a synchronous connection. The comparison gives us a reason to look for fit between the method and the work before standardising either approach.
Task fit also concerns the model. A peer-reviewed experiment with 758 Boston Consulting Group consultants found strong gains across 18 tasks chosen to sit within GPT-4's capabilities. AI users completed 12.2 per cent more tasks and worked 25.1 per cent faster on average. Their work also received higher quality scores. On one complex task chosen to sit outside the model's capability, people using AI were 19 percentage points less likely to reach the correct answer.
The BCG study tested only one task beyond the model's capability, so that figure has a narrow base. The broad point is still useful for work design: the value of AI changed within the same professional job. A prescribed workflow needs to say which part of the task the AI is expected to perform and what evidence supports that choice. A single route through a mixed task can place AI in a part of the work where its output needs a different form of human judgement.
People also adapt their work once AI becomes familiar. METR's 2025 randomised study of experienced open-source developers found that allowing early-2025 AI tools made 246 real coding tasks take 19 per cent longer. The 16 developers expected a speed gain before the study and still believed AI had made them faster afterwards.
When METR repeated the research with later tools, the researchers judged the new estimate unreliable. Developers who valued AI were less willing to take part in a trial that removed it for half their tasks. Between 30 and 50 per cent said they held back some tasks because they didn't want those tasks assigned to the no-AI condition. Running several agents at once also made time spent harder to record.
This evidence comes from software development and doesn't explain the Gap result. It shows how the act of measurement can lose contact with changing work. People select different tasks and overlap activities. Their beliefs about speed may also differ from observed time. A human-AI method needs an evaluation that can still describe what people actually do.
Treat the method as a design hypothesis
Teams often prescribe an AI workflow before they have observed the existing one. A few complete cases can reveal when people ask AI for help, what they retain and where a colleague enters. The observation should include the final check and any rework after the apparent finish. This gives the proposed method something concrete to improve.
The team then needs to name the problem the method is intended to solve. Weak briefs may call for a simple prompt aid. Loss of a colleague's knowledge may call for time together before drafting. These changes impose different work and should be tested separately where possible. Bundling them makes it difficult to learn which part helped.
The method's demands need their own measures. Count completion and total elapsed time through to an accepted result. Record later rework. Ask people where they left the prescribed route and what happened at that point. A repeated workaround gives evidence about the method and its setting; it shouldn't be recorded only as a worker failing to comply.
People also need a usable alternative when a dependency fails. Pairs unable to meet could add their views to a shared note and ask the AI to identify agreements and gaps. A time-sensitive case could allow one person to draft before a colleague reviews the result. These alternatives will have their own effects, which can be observed.
The organisation should decide in advance when to pause or revise the method. A large rise in non-completion is one possible trigger. Another is added coordination time with no improvement in the accepted work. In the Gap study, 37 per cent of the treatment pairs with usable behaviour data couldn't start the synchronous route. A local pilot showing a similar pattern would justify investigation before wider use.
Ruth Schmidt's recent work with Philip Cash and Weston Baxter gives us a useful way to review these choices. Their review of 12 behavioural design processes found that process guidance often mentions iteration without giving people enough help to decide when to start, stop, repeat or change it. They argue that a process should respond to its context and make the required practices and capabilities clearer.
Applied here, the review suggests that an AI workflow should remain provisional. A team can state the conditions it expects, including available time and a reliable way for colleagues to connect. It can also record intermediate results. Staff may learn when the AI needs a detailed brief, or discover that a shared discussion improves their own understanding even when the generated document is weak. Task performance remains the primary outcome. The intermediate results help explain what the organisation is building or eroding as people repeat the method.
Future experiments could randomise the working method within the same session and keep facilitation consistent. They could compare a strict sequence with an optional version and allow the time people would usually have. Researchers also need the interaction record: the human discussion, the material passed to the model, the model's first response and the revisions that followed. Independent human assessment and later use of the document would give a quality measure with less dependence on length.
A team can make the same comparison on a smaller scale. Describe the current route and its results. Introduce one designed change, then observe the route people and AI actually take. The difference between the designed and observed work becomes the material for the next design decision.
The research has not yet tested a lighter or more flexible version of Create-Out-Loud. It also leaves open how the method would perform during ordinary work, once people had learnt it and could choose a suitable moment to use it. Those comparisons are needed before this kind of protocol becomes standard practice.
Human-AI Performance
© 2026 Alterkind Ltd. All rights reserved.
Human-AI Performance™ is a proprietary methodology developed by BehaviourStudio using our Behaviour Thinking® framework. All content, tools, systems, and resources presented on this site are the exclusive intellectual property of Alterkind Ltd.
You’re welcome to use, share, and adapt these materials for personal learning and non-commercial team use.
For any commercial use, redistribution, or integration into client work, services, or paid products, please contact lauren@laurenakelly.com to discuss licensing terms.
Icons by Creative Mahira, The Noun Project.
Thanks to Nicholas Edell, Valentina Tan and multiple VPs implementing AI for your feedback during development.
LICENSE
Based on work by Lauren A Kelly.
For commercial licensing contact: lauren@laurenakelly.com
