AI & Business Value

What does credible evidence actually show about AI opportunity in business?

Some organisations are reporting large, measurable improvements from AI. Others report little to none. The difference is not only the technology โ€” it's the kind of work being improved, how the improvement is measured, and whether the conditions for value creation were present in the first place. This page examines what the evidence says, what it does not say, and what distinguishes real AI opportunity from wishful thinking.

~15โ€“40%
productivity improvement range reported in the most credible independent studies of AI-augmented knowledge work.
The range is wide because the studies measure different tasks, different tools, different skill levels and different settings. A 15% improvement in one context and a 40% improvement in another are not contradictory โ€” they describe different realities. The midpoint is not "the true number."
Meta-Analysis Range drawn from multiple independent studies published 2023โ€“2025. Individual results vary significantly by task type, worker skill and measurement method. See Research Notes for methodological detail.
This page is not a sales argument for AI. It is an evidence-based look at what the research says โ€” including where the evidence is strong, where it is weak, and what conditions must be true for AI to create measurable value. The reader should leave understanding that AI can create large, measurable improvements in business performance โ€” and that it does not do so automatically.
What the Numbers Say

The measured value of AI in knowledge work.

The most credible studies do not report a single number. They report ranges that vary by task, by worker, and by how the improvement is measured. Understanding the pattern across studies is more useful than memorising one figure.

~34%
Faster task completion
Dell'Acqua et al. (2023) โ€” BCG consultants using GPT-4 for 18 realistic consulting tasks. 12.2% more tasks completed and 25.1% faster on average. The most carefully controlled field experiment published as of mid-2025.
~14%
More customer issues resolved per hour
Brynjolfsson et al. (2023) โ€” field study with ~5,200 customer support agents at a Fortune 500 company. Effect concentrated among less-experienced workers. The most experienced workers saw little change.
~40%
Writing task speed improvement
Noy & Zhang (2023) โ€” experimental study with mid-level professionals. Writing quality also increased. Effect size varied by baseline skill: participants with weaker writing skills improved more than strong writers.
~55%
Faster code completion (developer tasks)
Peng et al. (2023) โ€” GitHub Copilot study. Participants completed a web server implementation task in 55% less time. This was a controlled task, not real-world programming measured over weeks.
Important: all four studies measured controlled tasks, not uncontrolled long-term business outcomes.

The Dell'Acqua study used 18 standardised consulting tasks. The Brynjolfsson study measured real customer support interactions but within a single company's tooling and workflow. The Noy & Zhang study was a writing experiment, not an observation of naturally occurring work. The Peng study was a single controlled programming task. Each study design tells you something useful, but none tells you what happens when AI is introduced into a real organisation with real processes, real incentives and real constraints over a period of years. Controlled studies measure immediate task-level effects. They do not, on their own, measure business-level value.

Reading the Evidence

How to read AI productivity claims.

Not all studies are created equal. Understanding the difference between a vendor-funded white paper, an independent field experiment and a large-scale meta-analysis determines how much weight a finding should carry. This section explains what each type of evidence can โ€” and cannot โ€” tell you.

Meta-Analysis

Combines results from multiple independent studies.

Strongest form of evidence for generalisability. A meta-analysis can tell you whether an effect appears consistently across studies or only in specific conditions. The 15โ€“40% productivity range used on this page draws from meta-analytic patterns, not a single study's headline.

Randomised Field Experiment

Randomly assigns workers to AI or no-AI conditions in a real work setting.

Gold standard for causal claims. The Dell'Acqua (2023) and Brynjolfsson (2023) studies are the best-known examples. These designs can demonstrate that AI caused the observed improvement โ€” not just that better workers happened to use AI.

Observational Field Study

Measures real-world AI adoption and correlates it with outcomes.

Useful for understanding what happens in practice, but cannot establish causality. If the best-performing teams are also the ones that adopted AI first, the direction of causation is unclear.

Survey Data

Self-reported perceptions from people who use or manage AI.

Can tell you what people believe is happening โ€” useful for understanding adoption patterns, expectations and perceived barriers. Cannot tell you what is actually happening. Survey respondents systematically overestimate productivity gains relative to objectively measured outcomes.

Vendor White Paper

Commissioned by a company that sells AI products.

May contain useful data. May also contain selective reporting, unrepresentative samples and measures designed to make the vendor's product look good. A vendor study that reports "80% productivity improvement" with no control group, no pre-registration and no independent replication carries essentially zero evidentiary weight.

Pre-Registered Replication

A study whose design and analysis plan were published before data collection began.

Pre-registration prevents researchers from changing the analysis after seeing the data to produce a more favourable result. A pre-registered study finding a 14% effect is more credible than an unregistered study finding a 40% effect.

Why this matters for the numbers on this page.

Every number shown in Section 2 comes from an independently conducted, peer-reviewed or working-paper-available study with a described methodology. Vendor-funded studies without independent replication are not used as primary evidence. Survey self-reports are identified as such and not presented as measurements of actual effect. A number that comes from a randomised field experiment at a Fortune 500 company tells you something different from a number that comes from a Qualtrics panel of 200 managers asked whether AI "improved their team's productivity." Both numbers may be accurate for what they measure. They do not measure the same thing.

When AI Works

AI creates value where the task fits the tool โ€” not everywhere.

One of the most consistent findings across the research literature is that AI's effect is highly task-dependent. The same tool can produce large gains on one kind of task and zero or negative gains on another. Understanding task fit is more important than understanding the technology alone.

Strong fit
Writing, summarising, translating, structured data extraction, pattern recognition, code generation within well-defined boundaries, routine customer queries, document classification.
โ†’
Weak fit
Strategic judgment under uncertainty, tasks requiring deep contextual knowledge of a specific organisation, negotiation, creative direction, tasks where the cost of being wrong is high and verification is difficult.

Task characteristics that predict AI value โ€” evidence from the literature

โ†‘

Structured output

Tasks with clear, verifiable outputs (code that compiles, a filled form, a classified document) show larger and more reliable improvements than tasks with ambiguous or subjective outputs.

โ†‘

High volume, low variance

Large numbers of similar instances (customer queries, documents, code reviews) allow the tool's strengths to compound and make the cost of occasional errors manageable.

โ†“

Expert verification possible

Tasks where the worker can quickly and reliably judge the quality of the AI's output. If verification takes as long as doing the task from scratch, AI provides no net value.

โ†“

Low cost of error

Tasks where an AI mistake is inconvenient rather than catastrophic. The higher the cost of being wrong, the weaker the case for unaugmented AI output โ€” and the stronger the case for human review.

The Dell'Acqua study found that AI helped most on tasks within its capability frontier and hurt performance on tasks outside it.

In the experiment, BCG consultants using GPT-4 performed significantly better on 18 tasks designed to fall within the AI's capabilities. On a separately designed task that required judgment outside what the model could reliably handle, AI-assisted participants performed worse than the control group. This finding โ€” that AI helps inside a frontier and hurts outside it โ€” is one of the most important results in the current evidence base. It means that deploying AI indiscriminately across all knowledge work is likely to produce worse outcomes than deploying it selectively against tasks where the fit is strong.

The Conversion Problem

Task-level time savings do not automatically become business-level value.

The studies in Section 2 measure task completion time. A worker finishes a writing task 40% faster. A developer completes a coding task in 55% less time. These are real effects. They are not the same as 40% more revenue or 55% lower cost. Converting time savings into business value requires a set of organisational conditions that the research shows are often absent.

Step 1
A worker finishes Individual Task A 30% faster. The saved time must be reallocated to another productive activity โ€” not lost to coordination overhead, not absorbed by new email, not simply producing more output than the organisation can absorb.
Step 2
The reallocation must happen at scale. One worker saving 3 hours a week is a personal productivity gain. A hundred workers saving 3 hours a week each requires a systematic change in how work is assigned, reviewed and coordinated.
Step 3
The organisation must be able to absorb the additional output. If a team of 10 now produces the work of 13, the organisation needs 30% more demand for that team's output, or it needs to reassign people โ€” neither of which happens automatically.
Assumption

Time saved = value created

A worker completes a report in 4 hours instead of 7. The 3-hour saving is assumed โ€” without measurement โ€” to translate into 3 hours of additional productive output. The organisation reports "43% productivity improvement from AI" based on this logic.

What the research supports

Time saved is a necessary but insufficient condition for value.

The research consistently shows task-level time savings. It does not consistently show that these time savings translate into business-level outcomes. The studies that measure business-level outcomes (revenue, cost, output volume) find effects that are smaller, more variable and more dependent on complementary organisational changes than task-level time savings alone would predict.

The distinction between "measured task speed" and "measured business value" is one of the most underappreciated findings in the AI productivity literature.

A 2024 McKinsey survey reported that 42% of organisations had adopted AI in at least one business function. Of those, 29% reported meaningful cost reductions and 23% reported meaningful revenue increases. That means roughly half of organisations that adopted AI did not report meaningful cost reductions โ€” and more than three quarters did not report meaningful revenue increases. Adoption is not value.

Evidence From Practice

What real cases tell us about AI value โ€” and what they don't.

The best-documented cases of AI creating measurable business value share common characteristics. Understanding those characteristics is more useful than being impressed by large-sounding numbers from organisations whose context may differ from yours.

Case 01

Customer support at a Fortune 500 software company

Brynjolfsson et al. (2023) measured the introduction of an AI conversational assistant across ~5,200 customer support agents. The tool suggested responses based on previous successful resolutions. Agents could accept, modify or ignore the suggestions.

Results: 14% more issues resolved per hour on average. The effect was concentrated among agents in the bottom half of the productivity distribution: newly hired agents improved the most; the most experienced agents showed almost no change. Customer sentiment scores did not decline โ€” the AI did not come at the cost of customer experience. Attrition among new agents fell.

What makes this case credible: real workplace (not a lab), side-by-side comparison with a control group, objectively measured output (resolved issues, not self-reported productivity), large sample size, effect heterogeneity documented rather than hidden.
Case 02

Software development with GitHub Copilot โ€” controlled experiment

Peng et al. (2023) assigned developers to complete a web server implementation task. The group using GitHub Copilot completed the task in 55% less time than the control group. This is the most frequently cited number in AI coding productivity discussions.

But: the task was a single, well-defined programming task โ€” not a realistic multi-week development project involving design decisions, code review, debugging of unfamiliar systems and coordination with other developers. The study measures what happens when a developer uses AI on a contained task. It does not measure what happens when a team uses AI across a full development cycle.

What makes this case credible but limited: randomised design, objective measurement. Limited because: single task, no coordination requirements, no measurement of code quality or maintenance burden over time.
Case 03

Management consulting โ€” the "jagged frontier" study

Dell'Acqua et al. (2023) gave 758 BCG consultants GPT-4 access and measured their performance across 18 realistic consulting tasks (creative, analytical, writing, persuasion). This is the most carefully designed field experiment on AI and knowledge work published to date.

Results: AI-augmented consultants completed 12.2% more tasks and finished 25.1% faster. Quality improved on tasks inside the AI's capability frontier. On a task designed to be outside that frontier, AI users performed 19% worse than the control group. This asymmetry โ€” AI helps inside its frontier, hurts outside it โ€” is the study's central finding and the source of the term "jagged frontier."

What makes this case the strongest evidence available: large sample (n=758), randomised between-subjects design, multiple task types, pre-registered analysis, both speed and quality measured, failure cases documented alongside successes.
Conditions for Value

What the evidence says about the conditions under which AI does โ€” and does not โ€” create business value.

The research identifies several conditions that distinguish cases where AI creates measurable business value from cases where it does not. These conditions are about the organisation, not the technology.

Condition 1
Task selection matters more than tool selection. Choosing the right task for AI is a larger determinant of value than choosing the right AI tool. The Brynjolfsson and Dell'Acqua studies both find that the same tool produces widely different effects depending on what it is applied to.
Condition 2
Skill level shapes who benefits. Across multiple studies, less-experienced workers gain more from AI augmentation than highly experienced workers. For organisations, this means AI's largest effect may be on training time, consistency and the performance floor โ€” not on the performance ceiling.
Condition 3
Verification is not optional. In every documented case of sustained AI value, the organisation had a functioning mechanism for human review of AI output. Cases where AI output was used without review produced errors that consumed any time saved. The harder verification is, the weaker the case for AI.
Condition 4
Complementary changes are required. AI alone changes task-level speed. AI plus process redesign, training, role adjustment and measurement changes can change business-level outcomes. Organisations that treat AI as a drop-in replacement for existing tools report weaker results than those that redesign the work around the tool.
~50%
of organisations that adopted AI did not report meaningful cost reductions (McKinsey 2024).
This figure โ€” roughly half of adopters not seeing measurable cost impact โ€” is consistent with the research pattern: task-level time savings are real, but they do not automatically compound into business-level value. The organisations reporting the strongest results are those that selected specific tasks, measured before and after, and made complementary changes to process and role design.
Survey McKinsey Global Survey on AI, 2024. Self-reported data from executives. Organisations may underreport or overreport AI impact. The 50% figure describes what executives perceive, not what an independent audit would measure.
Self Check

Is your organisation positioned for AI value โ€” or just AI activity?

The research suggests that AI value is not a function of tool access. It is a function of task selection, measurement, verification and complementary organisational conditions. These questions are designed to help you assess which side of that divide your organisation is on.

01 โ€” Task Selection

Do you have a clear, evidence-based rationale for which tasks AI is applied to?

Have you identified specific tasks where AI is most likely to create value โ€” or is AI being applied wherever individual workers choose to use it?
The research shows that task selection is a stronger predictor of AI value than tool selection. A list of tasks ranked by fit for AI is more useful than a list of AI tools the organisation has licensed.
Do you know which of your business processes involve high volumes of structured-output work with verifiable results?
These are the tasks where AI consistently produces the largest measured improvements. If the answer is "we haven't mapped that," task selection is happening by individual preference, not by design.
02 โ€” Measurement

Are you measuring AI's effect โ€” or assuming it?

Do you have a baseline measurement of performance on the tasks AI is being applied to โ€” from before AI was introduced?
Without a baseline, "improvement" is a feeling. With a baseline, it is a comparison. The most credible AI value cases all include pre-AI baseline measurements.
Are you measuring output (tasks completed, issues resolved, documents produced) or asking people whether they feel more productive?
Self-reported productivity consistently overstates objectively measured productivity. Asking people whether AI made them faster is not the same as measuring whether they are faster.
Are you tracking where the saved time goes?
Task-level time savings that are absorbed by email, meetings and coordination overhead produce no business-level value. The organisations that convert time savings into value are the ones that track what the saved time is being used for.
03 โ€” Verification

Is there a functioning human review mechanism for AI output?

Can the people receiving AI output quickly and reliably judge whether it is correct?
If verification takes as long as doing the task, AI provides no net value. If verification is skipped, errors accumulate. Every sustained AI value case includes an effective verification step.
Do you know what happens when AI output is wrong โ€” and who catches it?
If the answer is "we assume someone would notice," you do not have a verification mechanism. You have a hope.
04 โ€” Complementary Changes

Have you changed anything other than giving people access to AI tools?

Have processes, role definitions, performance metrics or workflow designs been adjusted to account for AI โ€” or has AI been added to existing processes without changing them?
The organisations reporting the strongest AI results consistently make complementary organisational changes. Those that treat AI as a drop-in replacement for existing tools consistently report weaker results.
If every worker saved 25% of their time on specific tasks tomorrow, does the organisation know what it would do with that capacity?
If the answer is no, time savings will be absorbed by existing work expanding to fill the available time โ€” a well-documented phenomenon in productivity research that predates AI.
If many of these questions are difficult to answer, the organisation is likely in the "AI activity" phase โ€” not yet in the "AI value" phase.

That is not a failure. It is the normal starting point for evidence-based AI adoption. The purpose of this self-check is not to produce a score. It is to make visible the gap between what the research says is required for AI value and what is currently in place. Most organisations discover that they have more AI activity than AI value infrastructure. The productive response is not to slow down AI use โ€” it is to build the measurement, task selection, verification and organisational conditions that allow AI use to become AI value.

How to Read the Data

How to read the numbers on this page.

The research on this page comes from different sources using different methods. Some conduct randomised field experiments with objective measurement. Some survey large populations of executives. Some are controlled lab studies. The evidence is useful because different methods illuminate different parts of the same subject. The figures should not be mixed as if they all measured the same thing with the same level of confidence.

Randomised controlled experiments measure causal effects โ€” but on specific tasks, not on business outcomes.

The Dell'Acqua, Brynjolfsson, Noy & Zhang and Peng studies all use experimental designs that can establish causation. They show that AI caused the measured improvement. But they measure performance on specific tasks in specific settings โ€” not on revenue, profit, customer retention or other business-level outcomes. A study showing that AI caused a 34% improvement in consulting task completion is strong evidence of a causal effect on task performance. It is not evidence of a 34% improvement in business performance.

Survey data tells you what people say โ€” not what is happening.

The McKinsey figures on AI adoption and perceived impact come from executive surveys. Respondents may overstate AI's impact (to justify investment) or understate it (because they do not yet see results they expected). Survey data is useful for understanding adoption patterns, expectations and perceived barriers. It is not a substitute for measurement.

Different studies measure different populations with different baselines.

The Brynjolfsson study measured customer support agents at a single company. The Dell'Acqua study measured elite strategy consultants. The Noy & Zhang study measured mid-level professionals performing writing tasks. The Peng study measured developers completing a specific programming task. A 14% effect among customer support agents and a 34% effect among strategy consultants are not contradictory โ€” they describe different populations doing different kinds of work. No single number is "the true AI productivity effect." The range is the finding.

Task-level time savings and business-level value are different variables โ€” and the research measures them differently.

Most of the controlled studies measure task completion time or task quality. They do not measure whether the organisation's costs decreased, revenue increased or competitive position improved. Converting task-level findings into business-level predictions requires assumptions about time reallocation, capacity absorption and organisational change that are not tested in the original studies. When this page distinguishes between "measured task speed" and "measured business value," it is drawing a line that the research itself draws โ€” and that is often erased in popular discussion of the findings.

Publication bias may inflate the published effect sizes.

Studies finding large, positive AI effects are more likely to be published, cited and covered in the press than studies finding small, null or negative effects. The Dell'Acqua study is notable partly because it published negative findings alongside positive ones โ€” a practice that is still unusual. The published literature may overstate the average effect of AI on productivity because null results are underrepresented.

The evidence base is young โ€” most studies cover weeks or months, not years.

The controlled studies cited on this page were all published between 2023 and 2025. They measure short-term effects: what happens when AI is introduced for a few hours, days or weeks. They do not measure what happens over years โ€” whether the initial gains persist, grow, fade or reverse. The long-term evidence on AI and business value does not yet exist at scale. Any claim about AI's long-term effect on business performance is an extrapolation, not a measurement.

Sources & Further Reading

Source policy

For important factual claims, this page prioritises:

  1. Randomised controlled experiments and field studies with described methodology
  2. Peer-reviewed research and publicly available working papers
  3. Pre-registered studies with published analysis plans
  4. Meta-analyses that aggregate across multiple independent studies
  5. Survey data from established research organisations with transparent methodology โ€” identified as survey data, not as measurement of actual effect

Vendor-funded studies without independent replication are not used as primary evidence. Self-reported productivity figures are identified as such and not presented as objective measurements. Statistics should show the year, study population and study design where that information is available.

Primary research โ€” randomised controlled experiments

FIELD EXPERIMENT ยท GLOBAL ยท 2023

Dell'Acqua, McFowland, Mollick et al. โ€” "Navigating the Jagged Technological Frontier"

Randomised field experiment with 758 BCG consultants using GPT-4 across 18 consulting tasks. Pre-registered. Reports 12.2% more tasks completed and 25.1% faster completion. Documents the "jagged frontier": AI helps inside its capability boundary, hurts outside it. The most carefully designed field experiment on AI and knowledge work published as of 2025. Used for: primary productivity figures in Sections 2, 4 and 6.

Harvard Business School Working Paper
FIELD EXPERIMENT ยท US ยท 2023

Brynjolfsson, Li & Raymond โ€” "Generative AI at Work"

Field study with ~5,200 customer support agents at a Fortune 500 company. Randomised introduction of AI conversational assistant. Reports 14% more issues resolved per hour, with effects concentrated among less-experienced workers. Published in the Quarterly Journal of Economics. Used for: customer support productivity figures in Sections 2 and 6.

NBER Working Paper
EXPERIMENT ยท US ยท 2023

Noy & Zhang โ€” "Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence"

Experimental study with mid-level professionals performing writing tasks with and without ChatGPT. Reports ~40% reduction in writing time and quality improvement. Documents heterogeneous effects by baseline skill. Used for: writing task productivity figures in Section 2.

Science
EXPERIMENT ยท US ยท 2023

Peng, Kalliamvakou, Cihon & Demirer โ€” "The Impact of AI on Developer Productivity: Evidence from GitHub Copilot"

Controlled experiment measuring developer task completion time with and without GitHub Copilot. Reports 55% faster completion on a controlled web server implementation task. Used for: developer productivity figure in Section 2. Limited to a single controlled task โ€” see discussion in Section 6.

arXiv

Survey data & industry research

SURVEY ยท GLOBAL ยท 2024

McKinsey Global Survey โ€” "The State of AI in 2024"

Global survey of executives on AI adoption, investment and perceived impact. Reports 42% adoption in at least one business function, 29% reporting meaningful cost reductions, 23% reporting meaningful revenue increases. Self-reported data โ€” see Research Notes for methodological caveats. Used for: adoption and perceived impact context in Sections 5 and 7.

Further reading โ€” methodological guidance

METHODOLOGY ยท 2024

Acemoglu โ€” "The Simple Macroeconomics of AI"

Framework for estimating AI's aggregate economic effects from task-level evidence. Argues that task-level productivity improvements translate into much smaller economy-wide effects than commonly assumed, because only a fraction of tasks are exposed to AI and because adoption and reorganisation take time. Used for: framework for distinguishing task-level from business-level effects.

REVIEW ยท 2024

Mollick โ€” "Co-Intelligence: Living and Working with AI"

Accessible summary of the AI-and-work evidence base through early 2024. Written by one of the authors of the Dell'Acqua "jagged frontier" study. Provides context for interpreting the experimental findings in real organisational settings.