What does credible evidence actually show about AI opportunity in business?
Some organisations are reporting large, measurable improvements from AI. Others report little to none. The difference is not only the technology โ it's the kind of work being improved, how the improvement is measured, and whether the conditions for value creation were present in the first place. This page examines what the evidence says, what it does not say, and what distinguishes real AI opportunity from wishful thinking.
The measured value of AI in knowledge work.
The most credible studies do not report a single number. They report ranges that vary by task, by worker, and by how the improvement is measured. Understanding the pattern across studies is more useful than memorising one figure.
The Dell'Acqua study used 18 standardised consulting tasks. The Brynjolfsson study measured real customer support interactions but within a single company's tooling and workflow. The Noy & Zhang study was a writing experiment, not an observation of naturally occurring work. The Peng study was a single controlled programming task. Each study design tells you something useful, but none tells you what happens when AI is introduced into a real organisation with real processes, real incentives and real constraints over a period of years. Controlled studies measure immediate task-level effects. They do not, on their own, measure business-level value.
How to read AI productivity claims.
Not all studies are created equal. Understanding the difference between a vendor-funded white paper, an independent field experiment and a large-scale meta-analysis determines how much weight a finding should carry. This section explains what each type of evidence can โ and cannot โ tell you.
Combines results from multiple independent studies.
Strongest form of evidence for generalisability. A meta-analysis can tell you whether an effect appears consistently across studies or only in specific conditions. The 15โ40% productivity range used on this page draws from meta-analytic patterns, not a single study's headline.
Randomly assigns workers to AI or no-AI conditions in a real work setting.
Gold standard for causal claims. The Dell'Acqua (2023) and Brynjolfsson (2023) studies are the best-known examples. These designs can demonstrate that AI caused the observed improvement โ not just that better workers happened to use AI.
Measures real-world AI adoption and correlates it with outcomes.
Useful for understanding what happens in practice, but cannot establish causality. If the best-performing teams are also the ones that adopted AI first, the direction of causation is unclear.
Self-reported perceptions from people who use or manage AI.
Can tell you what people believe is happening โ useful for understanding adoption patterns, expectations and perceived barriers. Cannot tell you what is actually happening. Survey respondents systematically overestimate productivity gains relative to objectively measured outcomes.
Commissioned by a company that sells AI products.
May contain useful data. May also contain selective reporting, unrepresentative samples and measures designed to make the vendor's product look good. A vendor study that reports "80% productivity improvement" with no control group, no pre-registration and no independent replication carries essentially zero evidentiary weight.
A study whose design and analysis plan were published before data collection began.
Pre-registration prevents researchers from changing the analysis after seeing the data to produce a more favourable result. A pre-registered study finding a 14% effect is more credible than an unregistered study finding a 40% effect.
Every number shown in Section 2 comes from an independently conducted, peer-reviewed or working-paper-available study with a described methodology. Vendor-funded studies without independent replication are not used as primary evidence. Survey self-reports are identified as such and not presented as measurements of actual effect. A number that comes from a randomised field experiment at a Fortune 500 company tells you something different from a number that comes from a Qualtrics panel of 200 managers asked whether AI "improved their team's productivity." Both numbers may be accurate for what they measure. They do not measure the same thing.
AI creates value where the task fits the tool โ not everywhere.
One of the most consistent findings across the research literature is that AI's effect is highly task-dependent. The same tool can produce large gains on one kind of task and zero or negative gains on another. Understanding task fit is more important than understanding the technology alone.
Task characteristics that predict AI value โ evidence from the literature
Structured output
Tasks with clear, verifiable outputs (code that compiles, a filled form, a classified document) show larger and more reliable improvements than tasks with ambiguous or subjective outputs.
High volume, low variance
Large numbers of similar instances (customer queries, documents, code reviews) allow the tool's strengths to compound and make the cost of occasional errors manageable.
Expert verification possible
Tasks where the worker can quickly and reliably judge the quality of the AI's output. If verification takes as long as doing the task from scratch, AI provides no net value.
Low cost of error
Tasks where an AI mistake is inconvenient rather than catastrophic. The higher the cost of being wrong, the weaker the case for unaugmented AI output โ and the stronger the case for human review.
In the experiment, BCG consultants using GPT-4 performed significantly better on 18 tasks designed to fall within the AI's capabilities. On a separately designed task that required judgment outside what the model could reliably handle, AI-assisted participants performed worse than the control group. This finding โ that AI helps inside a frontier and hurts outside it โ is one of the most important results in the current evidence base. It means that deploying AI indiscriminately across all knowledge work is likely to produce worse outcomes than deploying it selectively against tasks where the fit is strong.
Task-level time savings do not automatically become business-level value.
The studies in Section 2 measure task completion time. A worker finishes a writing task 40% faster. A developer completes a coding task in 55% less time. These are real effects. They are not the same as 40% more revenue or 55% lower cost. Converting time savings into business value requires a set of organisational conditions that the research shows are often absent.
Time saved = value created
A worker completes a report in 4 hours instead of 7. The 3-hour saving is assumed โ without measurement โ to translate into 3 hours of additional productive output. The organisation reports "43% productivity improvement from AI" based on this logic.
Time saved is a necessary but insufficient condition for value.
The research consistently shows task-level time savings. It does not consistently show that these time savings translate into business-level outcomes. The studies that measure business-level outcomes (revenue, cost, output volume) find effects that are smaller, more variable and more dependent on complementary organisational changes than task-level time savings alone would predict.
A 2024 McKinsey survey reported that 42% of organisations had adopted AI in at least one business function. Of those, 29% reported meaningful cost reductions and 23% reported meaningful revenue increases. That means roughly half of organisations that adopted AI did not report meaningful cost reductions โ and more than three quarters did not report meaningful revenue increases. Adoption is not value.
What real cases tell us about AI value โ and what they don't.
The best-documented cases of AI creating measurable business value share common characteristics. Understanding those characteristics is more useful than being impressed by large-sounding numbers from organisations whose context may differ from yours.
Customer support at a Fortune 500 software company
Brynjolfsson et al. (2023) measured the introduction of an AI conversational assistant across ~5,200 customer support agents. The tool suggested responses based on previous successful resolutions. Agents could accept, modify or ignore the suggestions.
Results: 14% more issues resolved per hour on average. The effect was concentrated among agents in the bottom half of the productivity distribution: newly hired agents improved the most; the most experienced agents showed almost no change. Customer sentiment scores did not decline โ the AI did not come at the cost of customer experience. Attrition among new agents fell.
Software development with GitHub Copilot โ controlled experiment
Peng et al. (2023) assigned developers to complete a web server implementation task. The group using GitHub Copilot completed the task in 55% less time than the control group. This is the most frequently cited number in AI coding productivity discussions.
But: the task was a single, well-defined programming task โ not a realistic multi-week development project involving design decisions, code review, debugging of unfamiliar systems and coordination with other developers. The study measures what happens when a developer uses AI on a contained task. It does not measure what happens when a team uses AI across a full development cycle.
Management consulting โ the "jagged frontier" study
Dell'Acqua et al. (2023) gave 758 BCG consultants GPT-4 access and measured their performance across 18 realistic consulting tasks (creative, analytical, writing, persuasion). This is the most carefully designed field experiment on AI and knowledge work published to date.
Results: AI-augmented consultants completed 12.2% more tasks and finished 25.1% faster. Quality improved on tasks inside the AI's capability frontier. On a task designed to be outside that frontier, AI users performed 19% worse than the control group. This asymmetry โ AI helps inside its frontier, hurts outside it โ is the study's central finding and the source of the term "jagged frontier."
What the evidence says about the conditions under which AI does โ and does not โ create business value.
The research identifies several conditions that distinguish cases where AI creates measurable business value from cases where it does not. These conditions are about the organisation, not the technology.
Is your organisation positioned for AI value โ or just AI activity?
The research suggests that AI value is not a function of tool access. It is a function of task selection, measurement, verification and complementary organisational conditions. These questions are designed to help you assess which side of that divide your organisation is on.
Do you have a clear, evidence-based rationale for which tasks AI is applied to?
- Have you identified specific tasks where AI is most likely to create value โ or is AI being applied wherever individual workers choose to use it?
- The research shows that task selection is a stronger predictor of AI value than tool selection. A list of tasks ranked by fit for AI is more useful than a list of AI tools the organisation has licensed.
- Do you know which of your business processes involve high volumes of structured-output work with verifiable results?
- These are the tasks where AI consistently produces the largest measured improvements. If the answer is "we haven't mapped that," task selection is happening by individual preference, not by design.
Are you measuring AI's effect โ or assuming it?
- Do you have a baseline measurement of performance on the tasks AI is being applied to โ from before AI was introduced?
- Without a baseline, "improvement" is a feeling. With a baseline, it is a comparison. The most credible AI value cases all include pre-AI baseline measurements.
- Are you measuring output (tasks completed, issues resolved, documents produced) or asking people whether they feel more productive?
- Self-reported productivity consistently overstates objectively measured productivity. Asking people whether AI made them faster is not the same as measuring whether they are faster.
- Are you tracking where the saved time goes?
- Task-level time savings that are absorbed by email, meetings and coordination overhead produce no business-level value. The organisations that convert time savings into value are the ones that track what the saved time is being used for.
Is there a functioning human review mechanism for AI output?
- Can the people receiving AI output quickly and reliably judge whether it is correct?
- If verification takes as long as doing the task, AI provides no net value. If verification is skipped, errors accumulate. Every sustained AI value case includes an effective verification step.
- Do you know what happens when AI output is wrong โ and who catches it?
- If the answer is "we assume someone would notice," you do not have a verification mechanism. You have a hope.
Have you changed anything other than giving people access to AI tools?
- Have processes, role definitions, performance metrics or workflow designs been adjusted to account for AI โ or has AI been added to existing processes without changing them?
- The organisations reporting the strongest AI results consistently make complementary organisational changes. Those that treat AI as a drop-in replacement for existing tools consistently report weaker results.
- If every worker saved 25% of their time on specific tasks tomorrow, does the organisation know what it would do with that capacity?
- If the answer is no, time savings will be absorbed by existing work expanding to fill the available time โ a well-documented phenomenon in productivity research that predates AI.
That is not a failure. It is the normal starting point for evidence-based AI adoption. The purpose of this self-check is not to produce a score. It is to make visible the gap between what the research says is required for AI value and what is currently in place. Most organisations discover that they have more AI activity than AI value infrastructure. The productive response is not to slow down AI use โ it is to build the measurement, task selection, verification and organisational conditions that allow AI use to become AI value.
How to read the numbers on this page.
The research on this page comes from different sources using different methods. Some conduct randomised field experiments with objective measurement. Some survey large populations of executives. Some are controlled lab studies. The evidence is useful because different methods illuminate different parts of the same subject. The figures should not be mixed as if they all measured the same thing with the same level of confidence.
Randomised controlled experiments measure causal effects โ but on specific tasks, not on business outcomes.
The Dell'Acqua, Brynjolfsson, Noy & Zhang and Peng studies all use experimental designs that can establish causation. They show that AI caused the measured improvement. But they measure performance on specific tasks in specific settings โ not on revenue, profit, customer retention or other business-level outcomes. A study showing that AI caused a 34% improvement in consulting task completion is strong evidence of a causal effect on task performance. It is not evidence of a 34% improvement in business performance.
Survey data tells you what people say โ not what is happening.
The McKinsey figures on AI adoption and perceived impact come from executive surveys. Respondents may overstate AI's impact (to justify investment) or understate it (because they do not yet see results they expected). Survey data is useful for understanding adoption patterns, expectations and perceived barriers. It is not a substitute for measurement.
Different studies measure different populations with different baselines.
The Brynjolfsson study measured customer support agents at a single company. The Dell'Acqua study measured elite strategy consultants. The Noy & Zhang study measured mid-level professionals performing writing tasks. The Peng study measured developers completing a specific programming task. A 14% effect among customer support agents and a 34% effect among strategy consultants are not contradictory โ they describe different populations doing different kinds of work. No single number is "the true AI productivity effect." The range is the finding.
Task-level time savings and business-level value are different variables โ and the research measures them differently.
Most of the controlled studies measure task completion time or task quality. They do not measure whether the organisation's costs decreased, revenue increased or competitive position improved. Converting task-level findings into business-level predictions requires assumptions about time reallocation, capacity absorption and organisational change that are not tested in the original studies. When this page distinguishes between "measured task speed" and "measured business value," it is drawing a line that the research itself draws โ and that is often erased in popular discussion of the findings.
Publication bias may inflate the published effect sizes.
Studies finding large, positive AI effects are more likely to be published, cited and covered in the press than studies finding small, null or negative effects. The Dell'Acqua study is notable partly because it published negative findings alongside positive ones โ a practice that is still unusual. The published literature may overstate the average effect of AI on productivity because null results are underrepresented.
The evidence base is young โ most studies cover weeks or months, not years.
The controlled studies cited on this page were all published between 2023 and 2025. They measure short-term effects: what happens when AI is introduced for a few hours, days or weeks. They do not measure what happens over years โ whether the initial gains persist, grow, fade or reverse. The long-term evidence on AI and business value does not yet exist at scale. Any claim about AI's long-term effect on business performance is an extrapolation, not a measurement.
Source policy
For important factual claims, this page prioritises:
- Randomised controlled experiments and field studies with described methodology
- Peer-reviewed research and publicly available working papers
- Pre-registered studies with published analysis plans
- Meta-analyses that aggregate across multiple independent studies
- Survey data from established research organisations with transparent methodology โ identified as survey data, not as measurement of actual effect
Vendor-funded studies without independent replication are not used as primary evidence. Self-reported productivity figures are identified as such and not presented as objective measurements. Statistics should show the year, study population and study design where that information is available.
Primary research โ randomised controlled experiments
Dell'Acqua, McFowland, Mollick et al. โ "Navigating the Jagged Technological Frontier"
Randomised field experiment with 758 BCG consultants using GPT-4 across 18 consulting tasks. Pre-registered. Reports 12.2% more tasks completed and 25.1% faster completion. Documents the "jagged frontier": AI helps inside its capability boundary, hurts outside it. The most carefully designed field experiment on AI and knowledge work published as of 2025. Used for: primary productivity figures in Sections 2, 4 and 6.
Harvard Business School Working PaperBrynjolfsson, Li & Raymond โ "Generative AI at Work"
Field study with ~5,200 customer support agents at a Fortune 500 company. Randomised introduction of AI conversational assistant. Reports 14% more issues resolved per hour, with effects concentrated among less-experienced workers. Published in the Quarterly Journal of Economics. Used for: customer support productivity figures in Sections 2 and 6.
NBER Working PaperNoy & Zhang โ "Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence"
Experimental study with mid-level professionals performing writing tasks with and without ChatGPT. Reports ~40% reduction in writing time and quality improvement. Documents heterogeneous effects by baseline skill. Used for: writing task productivity figures in Section 2.
SciencePeng, Kalliamvakou, Cihon & Demirer โ "The Impact of AI on Developer Productivity: Evidence from GitHub Copilot"
Controlled experiment measuring developer task completion time with and without GitHub Copilot. Reports 55% faster completion on a controlled web server implementation task. Used for: developer productivity figure in Section 2. Limited to a single controlled task โ see discussion in Section 6.
arXivSurvey data & industry research
McKinsey Global Survey โ "The State of AI in 2024"
Global survey of executives on AI adoption, investment and perceived impact. Reports 42% adoption in at least one business function, 29% reporting meaningful cost reductions, 23% reporting meaningful revenue increases. Self-reported data โ see Research Notes for methodological caveats. Used for: adoption and perceived impact context in Sections 5 and 7.
Further reading โ methodological guidance
Acemoglu โ "The Simple Macroeconomics of AI"
Framework for estimating AI's aggregate economic effects from task-level evidence. Argues that task-level productivity improvements translate into much smaller economy-wide effects than commonly assumed, because only a fraction of tasks are exposed to AI and because adoption and reorganisation take time. Used for: framework for distinguishing task-level from business-level effects.
Mollick โ "Co-Intelligence: Living and Working with AI"
Accessible summary of the AI-and-work evidence base through early 2024. Written by one of the authors of the Dell'Acqua "jagged frontier" study. Provides context for interpreting the experimental findings in real organisational settings.