AI Tech Observer

Artificial Intelligence

How AI Actually Improves Productivity at Work: By Cutting Waiting, Switching and Rework

AI is valuable not merely when it produces a first draft faster, but when it helps complete quality-bound work sooner without moving time into verification and rework.

By AI 科技观察19 min read
A knowledge worker organising scattered tasks into an orderly AI-assisted workflow
AI-generated editorial illustration: AI creates value by reducing waiting, switching and rework.Source:AI Tech Journal / GPT Image 2 · AI-generated image

A customer reply generated in two minutes does not mean that the customer received a usable answer two minutes later. It may still need supporting material, figures checked, terminology aligned, and a colleague to judge exceptional cases. If the answer is then returned for revision, the two minutes saved at the outset have merely created a small gap at the front of the process. Conversely, if the reply is accurate, compliant, clears review at the first pass and leaves no follow-up questions for the next colleague, writing it faster is a genuine productivity gain. The issue is not whether to acknowledge this. It is that the speed at which a first draft appears on screen is not evidence that the work as a whole was completed faster.

The “not typing time” in the title should be read as a measurement reminder, not a literal denial. AI can indeed shorten direct tasks. In a pre-registered experiment in professional writing, 444 professionals were randomly given access to ChatGPT; on their second task, they took about 10 minutes less, or roughly 37% less time, and their blind-rated scores were 0.45 standard deviations higher. [1] In another pre-registered experiment involving 758 management consultants, the AI group completed more of the 18 tasks judged in advance to be within GPT-4’s capability range, was 25.1% faster on average, and produced better-quality work. [2] Yet in METR’s randomised trial of developers familiar with mature open-source projects, allowing the use of early-2025 AI tools increased completion time by 19%. [3]

These three findings do not conflict. Together, they show that efficiency is neither model response time nor the number of additional pages, lines or emails generated today compared with yesterday. The more dependable unit is a complete piece of work: from the appearance of usable input to an outcome being accepted and delivered at the agreed quality, or being ready for direct use in the next step. In between are people waiting for people, people waiting for material, recovering their bearings after switching tools, checking the model’s claims, testing, revising, and downstream returns. Only when this complete cycle is shorter, quality has not fallen, and subsequent costs have not been quietly shifted elsewhere is it reasonable to say that AI has improved work efficiency.

That is also why waiting, switching and rework should not be treated as three AI effects already established everywhere. Verification and rework have direct but limited evidence in particular tasks; waiting and context switching are better treated as hypotheses that each workflow must test. The useful question is not “should we use AI?” but: in this specific workflow, where does AI take time away, and where does it put time instead?

Contents

  1. 1.1. The Unit of Efficiency Is a Complete Work Cycle That Meets the Quality Threshold
  2. 2.2. Whether Waiting and Handovers Are Shorter Must Be Measured in the Specific Workflow
  3. 3.3. The Net Cost of Context Switching Must Be Tested in Comparable Tasks
  4. 4.4. Verification and Rework Determine Whether a Faster First Draft Produces a Net Gain
  5. 5.5. Teams Should Assess an AI Workflow Through Its Complete Cycle, Quality and Rework
  6. 6.Sources and citations

1. The Unit of Efficiency Is a Complete Work Cycle That Meets the Quality Threshold

Calling “writing faster” efficient is tempting because writing is the easiest part to see. Open a tool, enter a request and watch a passage appear: the timer immediately produces an attractive number. But the end of most work is not a first draft; it is a quality gate being passed. A product manager receives a conclusion that can be acted on; a customer receives an answer that resolves the problem; code passes testing and can be merged; an operational rule is correctly executed in a system. If the last step still needs extensive evidencing, revision or back-and-forth confirmation, generation speed at the front of the workflow does not represent final output.

This does not require treating every direct acceleration as false efficiency. The Noy and Zhang experiment matters precisely because it measured both time and quality rather than merely asking participants whether they felt faster. Among the 444 professionals randomly given the tool, the second professional-writing task took about 10 minutes less, a reduction of about 37%, while blind-rated quality rose by an average of 0.45 standard deviations. The result shows that, in short writing tasks with clear boundaries and quality assessable by blind review, faster completion and a better result can occur together. [1] The management-consultant experiment by Dell’Acqua and colleagues offers a similar but distinct setting: across 18 tasks pre-judged to be within the model’s capability range, the AI group completed 12.2% more tasks, was 25.1% faster on average, and showed a significant improvement in quality. [2] The SPACE framework for developer productivity is therefore particularly useful: productivity cannot be represented by one metric, one dimension or one activity. [4]

The conclusion these two experiments support is specific. Where task boundaries, quality standards and methods of assessment are all relatively clear, AI’s direct acceleration can be a real gain. Their limits matter just as much: a writing task completed faster does not mean a team’s project, customer response or delivery cycle will necessarily be shorter. Noy and Zhang’s task was conducted online and lasted around 20 to 30 minutes; it did not include long-term collaboration or downstream maintenance. The consultant experiment’s in-range tasks were screened in advance, which is not the same as putting AI into all everyday work and obtaining the same results. [1] [2]

When measuring a workflow, then, do not begin by asking how many words the model wrote. Specify the output and the quality threshold. For a sales-operations team, the output might be “from receiving a customer request to a sendable, traceable reply”; for an engineering team, “from confirming a task to a change that passes testing and can be safely released”; for a product team, perhaps “from research input to decision material adopted by the responsible lead”. The same input, with a different end point, produces a wholly different timing result. Time spent on the first draft remains worth recording, but it is one entry in the ledger, not the ledger itself.

This definition has another easily missed advantage: it exposes cost shifting. If AI lets a front-line worker complete a record faster but hands the burden of checking it to the back office; if it lets an engineer write a change faster but leaves testing and debugging to consume more time; if it makes a proposal look more complete but leaves the decision-maker longer to identify unreliable evidence, efficiency has not appeared from nowhere. The timer has simply been placed earlier in the process. Measuring the complete cycle does not reject speed; it prevents local speed from being mistaken for the speed of the whole system.

A quality threshold need not be a vague satisfied-or-unsatisfied survey. It may be whether an answer is correct, whether its evidence is traceable, whether its format can enter a downstream system, whether rules were followed, or whether a change survives testing. Different kinds of work combine these conditions differently, so no single number of minutes can judge them all. What must remain stable are the conditions of comparison: the end point for the same kind of task, the standard used by the person accepting it, and the rule that a return still belongs to the original work item. Only then will AI’s apparent speed not be improved by lowering the standard, narrowing the definition of delivery, or excluding difficult cases from the sample.

This also explains why output volume is often a seductive distraction. Ten first drafts, ten suggestions or ten code snippets can make work seem to have advanced tenfold. But if most are not adopted, or must be individually repaired downstream, they are closer to work-in-progress inventory than completed output. By contrast, a shorter document that the next person can use immediately is often closer to efficiency than a pile of text awaiting selection. The complete-cycle perspective does not demand slower work; it puts usability back into the definition of speed.

2. Whether Waiting and Handovers Are Shorter Must Be Measured in the Specific Workflow

Waiting at work rarely appears as a single button. It may be the pause while material is incomplete, an hour waiting for a specialist colleague to reply, a day in which a case sits in a queue without anyone taking it on, or the gap after one person has finished but before the next knows they can begin. To users, these periods are often more frustrating than writing a paragraph. But “AI reduces waiting” cannot be established by intuition alone. It is necessary to distinguish between AI making one local action faster and work actually reaching the next person able to handle it sooner.

Some field evidence can suggest a direction, but cannot substitute for that measurement. Brynjolfsson, Li and Raymond studied one company’s 5,172 customer-service workers. After generative AI was deployed, the average number of issues resolved per hour rose by 15%; the proportion of customers asking for a manager fell by nearly 25% relative to a baseline of about 6%. But the researchers explicitly had no records of actual escalations: the measure was a customer request to escalate, not an actual transfer of the customer, nor the waiting time for a handover. [5] In a six-month randomised field experiment across 66 firms and 7,137 knowledge workers, Dillon and colleagues also found that the 80% of the treatment group who used the tool in the latter part of the experiment spent about two fewer hours a week on email and worked less outside normal hours. The study did not, however, detect a change in the number or composition of tasks. [6]

The two studies concern phenomena at different levels. The first shows that, in a particular customer-service setting, the resolution rate rose and requests for managerial help fell. The second shows that an individual’s email time can be compressed without an automatically detectable reorganisation of tasks. These are local changes worth observing; they cannot be translated into “the queue has shortened”, “human handovers have fallen” or “organisational throughput has increased”. The first study lacks data on actual escalations; the second did not measure waiting, collaborative quality, revenue or customer value. [5] [6]

Waiting is therefore a workflow hypothesis to test, not a promotional claim. Consider a customer process that requires analysis, review and a reply. If, by the time the analyst has finished, AI has organised the sources, preliminary classification and points requiring confirmation so that the reviewer can use them directly, the median time from the previous step being ready to hand over to the next person starting effective work may fall. Actual escalations, follow-up questions and returns caused by incomplete material may also fall. “May” is crucial here. Only by recording the true readiness time, actual start time and actual handover events for comparable tasks can a team know whether the change occurred.

Measurement should also separate different waits. Waiting for material, for a colleague’s reply, for review, for an external system and for model generation are not the same cost, and the same change will not necessarily address them. Treating a customer’s request for a manager or a shorter period spent on email as a substitute for every form of waiting makes the real bottleneck easier to miss. If the first draft is faster but the review queue does not move, the information needed for handover remains incomplete, or the model itself adds waiting, this workflow cannot claim to have saved the waiting named in the title merely because its front end is faster.

Waiting also magnifies small gaps in collaboration. If the preceding colleague has not left the sources, assumptions and open questions, the next person may need time to decide whether the material is usable even if it arrives immediately. Conversely, a summary that clearly marks uncertainty may not shorten model generation, but it may shorten the recipient’s decision time. Both changes are worth observing, but they are not the same metric. Conflating handover design with queue management leaves improvement at the level of “it feels smoother” and makes it impossible to identify which step truly reduced idle waiting.

When testing a waiting hypothesis, then, it is best not to treat the AI condition as a one-off magical intervention. Specify where it sits in the workflow: whether it helps organise material at the input stage, offers preliminary classification during processing, or generates an evidence-backed summary at handover. Each placement has different potential gains and different ways to fail. If it only lets an individual write faster without changing whether the next person can begin immediately, waiting has no reason to fall automatically. Only if it makes information usable earlier is it worth continuing to observe queue and handover timestamps.

3. The Net Cost of Context Switching Must Be Tested in Comparable Tasks

“Switching between fewer applications” sounds like an obvious efficiency goal, but in practice it is more complicated. To answer a customer question, someone may need to move between a ticketing system, a knowledge base, chat records, a spreadsheet and internal systems. An engineer may move between code, tests, documentation and monitoring; a product manager may also need to consult research, raw data and colleagues’ views. A large number of applications does not automatically mean wasted work. The key is not how often the screen changes, but whether each switch means finding evidence again, recovering task state, transcribing information or reinterpreting what came before.

Interviews by Jahanlou and colleagues with 15 knowledge workers from five product teams found that switching software while completing the same task brings data conversion or transfer, refocusing and extra cognitive load. At the same time, proprietary functions, collaboration, policy and privacy requirements can make a multi-tool workflow necessary. [7] Larger-scale behavioural data must not be read too simply either. A preprint under review by Vaid and Whillans recorded 103 million application events, second by second, among 1,017 workers in eight Indian organisations. On days when AI was used, total switching was higher; in the short window after each use of AI, however, fewer applications and switches occurred, people stayed longer in a single application, and activity was more predictable. The authors also make clear that such records cannot observe attention, intent, quality or end-to-end output, and cannot tell whether AI caused the reorganisation or whether different kinds of workdays were already more likely to involve AI use. [8]

This is a useful reminder to distinguish two statements. “Short-term activity around AI is more concentrated” is an observation; “AI has reduced the cost of context switching and increased output” is not. Interviews can explain why switching is tiring and why some switching is unavoidable. Application logs can show an association in behavioural patterns, but cannot show that tasks were completed better. Multiple tools sometimes preserve specialised capability, trustworthy sources or necessary collaboration. Removing those steps in pursuit of a single entry point may instead hide quality costs. [7] [8]

It is therefore more useful to state a context hypothesis. In a fixed task, if AI can organise traceable material, constraints and unresolved questions relevant to the current action for the user, the time needed to restore task state, repeated searches and unproductive back-and-forth between applications may fall without quality falling. It still needs comparison before it can hold. The comparison should be between tasks of similar complexity and involving similar roles, not between a simple AI day and an unusually hectic non-AI day. The record should not contain only the number of applications; it should include the purpose of each switch: evidence gathering, data entry, enquiry, approval, verification, or a renewed search because the previous answer was unreliable.

This perspective changes how improvement is pursued. Rather than broadly telling staff to “open fewer windows”, find the moments when material is repeatedly passed along, background repeatedly explained, or constraints have to be recalled from scratch. A good supporting step may carry the question, sources and open issues into the next step, reducing the labour of reconstructing context. A bad supporting step may add another chat window and create more summaries without sources, increasing the work of checking them. Whether there is a net gain still depends on the complete cycle and quality outcome for comparable tasks, not on whether the interface looks tidier.

4. Verification and Rework Determine Whether a Faster First Draft Produces a Net Gain

The most common misunderstanding of generative AI is to treat “producing a first version” as “completing a first version”. In work involving facts, implicit rules, project history or boundary conditions, a model’s apparently fluent answer may only mean that the centre of human work has shifted from writing to judgement: which material can be adopted, which sources must be traced, which exceptions cannot be missed, and what needs testing. This shift is not necessarily bad. If checking is brief and the rules are clear, faster first drafts may still win out. If checking becomes line-by-line auditing, repeated trial and error, and downstream clean-up, the initial gain will be consumed.

METR’s randomised controlled trial makes this offset concrete. It asked 16 developers familiar with mature open-source projects to handle 246 pre-defined real tasks. When early-2025 AI tools were allowed, completion time rose by 19%. In the sub-sample with valid annotated video, about 9% of time went on reviewing or cleaning up AI output, about 4% on waiting for the model, and less than 44% of AI-generated content was accepted. This does not mean AI must slow every kind of programming work. It shows that, in this small sample of maintainers familiar with their projects and using a fixed snapshot of tools, prompting, waiting, review and clean-up can outweigh the benefit of direct generation. [3] The consultant experiment by Dell’Acqua and colleagues also shows that capability boundaries cannot be ignored. In one complex management task selected as outside the model’s capability range, the probability of producing the correct solution was, on average, 19 percentage points lower under the AI conditions—not “19% lower” in relative terms. In the detailed results, the control group was about 84.5%, while the two AI conditions were 60% and 70.6%. [2] A human-factors review groups such effects into several possible mechanisms: work shifts from production to evaluation, workflows are adversely reorganised, systems interrupt users, and simple tasks become simpler while difficult tasks become more difficult. [9]

The value of this evidence is that it turns “rework” from an abstract concern into a measurable component. Its limits must also remain intact. METR measured task-implementation time for particular developers in particular mature projects with a snapshot of early-2025 tools; its sample was small and it did not cover the complete lifecycle after deployment. The consultant study’s out-of-range result came from one complex task, so it cannot estimate a general failure rate across all work. [3] [2] These studies cannot establish that “AI always creates rework”, just as the positive experiments above cannot establish that AI errors can always be readily absorbed.

A more accurate approach is to set a verification budget for a workflow. It includes writing prompts or preparing inputs, waiting for the model, human review, source checking, testing, revision, returns and downstream correction. The real gain is not how much content the model produced first, but how much less total time was spent before the quality threshold. Subtasks with clear rules, automatic checks and complete inputs are more likely to turn faster first versions into net gains. Tasks involving implicit context, rare exceptions or high-risk judgements require actual data before a conclusion is drawn. “More likely” here is a screening hypothesis, not a prior commitment to any one tool or occupation.

First-pass approval rate is a particularly useful signal, but it cannot stand alone. If that rate rises while human verification time doubles, there may still be no saving. If the number of returns falls while later customer correction increases, the quality gate has been set too narrowly. Only by putting first-pass approval, human verification, returns and downstream correction into the same cycle can a team see whether it has reduced rework or merely shifted it to unseen people and later stages.

5. Teams Should Assess an AI Workflow Through Its Complete Cycle, Quality and Rework

Taken together, these principles mean that a team should not be trying to answer “does AI generally improve productivity?” The question is too large, and it readily invites a search for a flattering answer in output volume, time logged in, or individual rankings. A more workable question is whether a clearly bounded, recurrent workflow can be completed faster with the same or better quality after AI is introduced. Where do waiting, context reconstruction, verification and rework occur? The answer may vary by task, role, tool version, input quality and way of collaborating. That does not weaken the value of measurement; it explains why measurement must sit inside the workflow.

DORA’s software-delivery guidance provides a useful point of reference. It advises looking not only at throughput but also at instability, including change lead time, deployment frequency, time to restore service, change failure rate and deployment rework rate. It also warns against allowing a single metric to dominate a complex system or comparing outside context. [10] SPACE makes the same point: developer productivity is not a single activity metric. [4] Dillon and colleagues’ experiment, meanwhile, suggests that spending less time on email as an individual does not automatically appear as a change in the number or composition of tasks. [6] None of these is an AI programme proven for every function; DORA is not a cross-functional causal study of AI. They offer a measurement discipline that does not treat local speed as the final answer.

In practice, the smallest test should begin with a repeatable workflow that has a clear acceptance criterion, not with a company-wide rollout. First fix the output and quality gate. For example, make “customer question received to responsible lead accepting a sendable reply” one work unit, or “requirement confirmed to change passing safety testing” another. Define in advance what counts as first-pass approval, what counts as a return, and which downstream errors still belong to this work. Without those boundaries, every comparison of speed becomes a different person’s interpretation of “complete”.

Next, open up the timeline rather than recording only one total duration. Record separate timestamps for the start of processing, inputs becoming ready, handover to the next person, that person’s actual start, model response, human verification and final acceptance. This makes the waiting hypothesis testable: is the change in waiting for inputs, people, review or the model? Record context reconstruction too. Was a particular back-and-forth switch needed to obtain a trustworthy source, or did the previous summary omit a constraint? Was a repeated search a necessary check, or was the same material repeatedly transcribed? Such records need not turn every person’s mouse trail into an assessment tool. Their purpose is to distinguish necessary labour from removable friction.

Finally, place the AI condition alongside a baseline comparable in complexity, role and task type. Compare median complete-cycle time, first-pass approval rate, human verification time, number of returns and downstream correction—not only first-draft speed. If the complete cycle shortens, quality is stable or improves, and rework has not been moved later, that can be treated as a genuine gain in this workflow. If only the first version is shorter, waiting is unchanged, or verification and returns increase, it should honestly be called local acceleration, shifted cost or an unverified change, not a broad improvement in efficiency. DORA’s warning against metric misuse applies here as well: once a number becomes an individual quota, people will tend to select simple tasks, split tasks or hide rework, and measurement itself will damage judgement. [10]

This comparison need not wait for perfect data. A small, time-limited comparison can be enough to reveal whether a workflow has an obvious cost shift, provided the results are interpreted honestly. If the sample is small or task differences are large, the conclusion should remain “what was observed in this set of tasks”, not become a commitment for the whole company. Writing uncertainty into the conclusion does not weaken improvement. On the contrary, it lets the next round measure only the waiting, context or rework still not understood, rather than expanding use on the strength of one attractive average.

This starts more slowly than “have AI write a little more”, but it identifies more quickly what is worth scaling. AI’s real potential value may be that quality-compliant direct tasks finish faster; it may also be that, across a complete work cycle, work spends less time stuck waiting, less time rebuilding context, and less time repeatedly reworking wrong answers. Keyboard speed is not enough. These elements must be counted with the quality threshold, observed one by one, tested in comparable tasks, and checked to ensure that costs have not reappeared at the next stage of the process. Acceleration in a direct task can be a genuine efficiency gain; efficiency across a complete workflow still has to be demonstrated by net changes in the complete cycle, waiting, context reconstruction, verification and rework.

Sources and citations

  1. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality

    Fabrizio Dell’Acqua, Edward McFowland III, Ethan Mollick, Hila Lifshitz, Katherine C. Kellogg, Saran Rajendran, Lisa Krayer, François Candelon, and Karim R. Lakhani · Organization Science · Abstract; §6; §6; §4.2; Figure 5; Table 7

  2. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity

    Joel Becker, Nate Rush, Beth Barnes, and David Rein · Model Evaluation & Threat Research · Abstract; §§3.2, C.1.4, C.2.8; pp. 1–3

  3. The SPACE of Developer Productivity: There’s More to It Than You Think

    Nicole Forsgren, Margaret-Anne Storey, Chandra Maddila, Tom Zimmermann, Brian Houck, and Jenna Butler · ACM Queue · Abstract

  4. Generative AI at Work

    Erik Brynjolfsson, Danielle Li, and Lindsey R. Raymond · The Quarterly Journal of Economics · Abstract; §VII.B; Table IV; §VIII

  5. Shifting Work Patterns with Generative AI

    Eleanor W. Dillon, Sonia Jaffe, Nicole Immorlica, and Christopher T. Stanton · American Economic Review: Insights · Abstract

  6. Task-Centric Application Switching: How and Why Knowledge Workers Switch Software Applications for a Single Task

    Amir Jahanlou, Jo Vermeulen, Tovi Grossman, Parmit K. Chilana, George Fitzmaurice, and Justin Matejka · Graphics Interface Conference · Abstract; pp. 1–2; Abstract

  7. Digital Fragmentation and Generative AI Use Across 103 Million Application Events

    Sumer S. Vaid and Ashley V. Whillans · arXiv · pp. 1, 8–11; pp. 10–11

  8. Ironies of Generative AI: Understanding and Mitigating Productivity Loss in Human-AI Interaction

    Auste Simkute, Lev Tankelevitch, Viktor Kewenig, Ava Elizabeth Scott, Abigail Sellen, and Sean Rintel · International Journal of Human–Computer Interaction · Abstract

  9. DORA’s Software Delivery Performance Metrics

    DORA · Google Cloud · “Throughput and instability”; “Common pitfalls”; “Common pitfalls”