Evidence

What the research says

Published studies of generative AI at work, including Statistics Canada's survey of Canadian businesses, each shown with its sample, its task and its caveat. None of these figures is an iSystematic result.

The rule

iSystematic publishes figures from published studies, never its own results. Each figure shows what was measured, on whom, in which task, and the caveat that has to sit beside it.

At least one counter-finding is always shown beside the gains. These figures describe generative AI in studies; they say nothing about the effect of an iSystematic workflow, solution or pack.

Every figure links to its original source. The register is rechecked against the originals each quarter, and a new study is added only after the original has been read.

Where studies found gains

Four studies that measured time, output or quality with generative AI, from controlled experiments to a survey of US workers.

40% less time
Average time on short professional writing tasks fell 40% when participants used ChatGPT, and evaluator-graded quality rose 18%.
453 college-educated professionals (marketers, grant writers, consultants, data analysts, HR professionals and managers) in a preregistered online experiment; half were randomly given ChatGPT. The tasks were paid, occupation-specific writing assignments of 20 to 30 minutes: press releases, short reports, analysis plans and delicate emails.
Short, self-contained tasks in an experiment, without context-specific knowledge; the authors say this may inflate the estimates. Quality was graded by evaluators.
Noy and Zhang, Science, 2023 ↗
15% more
Customer-support agents with a generative AI assistant resolved 15% more issues per hour on average; less experienced and lower-skilled agents improved in both speed and quality.
5,172 customer-support agents at one firm, a Fortune 500 company that sells business-process software. The assistant was introduced in stages, not in a randomised trial. Task: customer-support chat, measured as issues resolved per hour.
The most experienced and highest-skilled agents saw small gains in speed and small declines in quality. The authors say the findings apply to one AI tool, used in one firm, in one occupation.
Brynjolfsson, Li and Raymond, Quarterly Journal of Economics, 2025 ↗
25.1% faster
Consultants using GPT-4 on tasks inside the AI's capability finished 25.1% more quickly, completed 12.2% more tasks, and produced work rated more than 40% higher in quality than a control group.
758 Boston Consulting Group consultants in total, in a preregistered experiment, randomly assigned to no AI, GPT-4, or GPT-4 with a prompt-engineering overview. These figures come from the 385 who worked on 18 realistic consulting tasks inside the AI's capability, such as developing new product ideas.
The gains held only for tasks inside the AI's capability. On a task outside it, consultants using AI were less likely to be correct (E3b). One firm, GPT-4 as it was in 2023, and a working paper.
Dell'Acqua et al., Harvard Business School working paper, 2023 ↗
5.4% of work hours
Generative AI users reported saving 5.4% of their work hours in the previous week; across all workers, including non-users, the saving came to 1.4% of total hours.
Real-Time Population Survey, November 2024 wave: 5,329 US adults aged 18 to 64 in an online panel, targeted and weighted to the Current Population Survey; the authors say it is not a random sample. The time-savings question went only to generative AI users (933 employed users, per the working paper).
Self-reported, from one week's recall: the extra hours a user said they would have needed without generative AI. The time saved was not measured directly.
Bick, Blandin and Deming, Federal Reserve Bank of St. Louis, 2025 ↗

Published studies, not our results. Ours will come from documented pilots.

Where studies found the opposite

Two results that cut against the gains: experienced developers who were slower with AI than they believed, and consultants who were less often right when the task lay outside the AI's capability.

19% slower
Experienced open-source developers took 19% longer to complete tasks when AI was allowed, though they estimated afterwards that AI had cut their completion time by 20%.
16 experienced developers, with about five years on their projects, working 246 real issues in their own mature repositories; each issue was randomly assigned to AI allowed or not allowed. The main tools were Cursor Pro with Claude 3.5 and 3.7 Sonnet, February to June 2025.
Early-2025 tools, and expert users on familiar, complex work. METR says the result does not show that AI fails to speed up most developers. In February 2026 METR gave the 95% confidence interval as 2% to 39% slower, said developers are likely more sped up now, and said its newer data give an unreliable signal because developers opted out of working without AI.
METR (Becker, Rush, Barnes and Rein), 2025 ↗
19 points less correct
Consultants using GPT-4 on a business problem outside the AI's capability were 19 percentage points less likely to reach a correct solution than consultants working without AI.
373 of the same 758 consultants, on a business problem-solving task using data and interviews, with the same preregistered random assignment.
The group without AI was correct about 84.5% of the time, against 60% and 70% in the two AI conditions. One firm, GPT-4 as it was in 2023, and a working paper.
Dell'Acqua et al., Harvard Business School working paper, 2023 ↗

Published studies, not our results. Ours will come from documented pilots.

In healthcare

One randomised trial of ambient AI scribes, which draft clinical notes. It measured physicians, not front desks.

About 9.5% less
Physicians spent about 9.5% less time writing each note with one of two ambient AI scribe tools (Nabla); the other tool (DAX) showed no significant change.
238 outpatient physicians across 14 specialties, randomised 1:1:1 to DAX, Nabla or no scribe; 72,369 visits over two months at one academic institution, measured as time-in-note per note in the Epic health record.
Single centre; physicians, not front desks; English-only visits. Burnout and workload were secondary measures and were not statistically tested; the authors say any improvement needs confirmation in larger trials. Clinicians reported that notes occasionally contained clinically significant inaccuracies. The scribes were used in only 29.5% (Nabla) and 33.5% (DAX) of visits, and time-in-note leaves out editing inside the vendors' apps.
Lukac et al., NEJM AI, 2025 ↗

Published studies, not our results. Ours will come from documented pilots.

In software engineering

One controlled experiment on a single coding task. It is not representative of small-business work, and iSystematic cites it only for enterprise engineering work.

55.8% faster
Developers with GitHub Copilot completed one coding task, writing an HTTP server in JavaScript, 55.8% faster than developers without it.
95 professional developers recruited on Upwork and randomised: 45 with Copilot, 50 without. Time ran until the code passed all 12 tests.
A single coding task; code quality was not measured, and time was measured only for those who finished. The 95% confidence interval ran from 21% to 89%. A preprint, not peer-reviewed. The authors work at Microsoft Research and GitHub, which make Copilot, and at MIT Sloan. Not representative of small-business work: iSystematic cites it only for enterprise engineering work.
Peng, Kalliamvakou, Cihon and Demirer, arXiv preprint, 2023 ↗

Published studies, not our results. Ours will come from documented pilots.

Adoption in Canada

Statistics Canada's survey of how many Canadian businesses use AI, and what the businesses that use it say it changed.

12.2%
12.2% of Canadian businesses reported having used AI to produce goods or deliver services in the previous 12 months, up from 6.1% in the second quarter of 2024.
Canadian Survey on Business Conditions, second quarter of 2025: a stratified random sample of 21,357 business establishments with employees; 9,103 responses, with calibrated weights; collected 1 April to 5 May 2025.
Use in producing goods or delivering services, not any use. It varied widely by industry: 35.6% in information and cultural industries, 31.7% in professional, scientific and technical services, 30.6% in finance and insurance, and 1.5% in accommodation and food services. Businesses with employees only.
Statistics Canada, 2025 ↗
47.2%
47.2% of Canadian businesses that used AI said it reduced employees' tasks to a small extent, and 89.4% reported no change in total employment.
Same survey as E8. The base is businesses that used AI (12.2% of all businesses), not all businesses.
Self-reported, over a 12-month horizon. Tasks were reduced to a large extent for 5.3%, a moderate extent for 32.4%, a small extent for 47.2%, and not at all for 15.1%. Employment rose at 4.3%, fell at 6.3% and was unchanged at 89.4%.
Statistics Canada, 2025 ↗

Published studies, not our results. Ours will come from documented pilots.

How to read a figure

A figure from a study is only as useful as the three things printed beside it.

01

The sample

Who was measured, how many, and how they were chosen. A randomised trial, a staggered rollout and a self-reported survey carry different weight.

02

The task

What people were doing when they were measured. A result on a short writing task says little about long, unfamiliar work, and a coding experiment says little about a front desk.

03

The caveat

What limits the result, as the source itself says: one firm, one tool, self-reported time, a preprint. Read it before quoting the number.

How iSystematic's own figures will come

iSystematic's own figures will come only from documented pilots, measured before and after: task time including human review, draft acceptance, corrections, escalations, duplicates, failures and cost. This page shows none yet.

We will publish what the measurement shows, including no change or a worse result, and never select only the good results. Each result will state the number of organisations, the task and the period, and will sit beside the published studies, never merged with them. A result is published only with the client's permission.

Start

Ask how a pilot is measured