Skip to main content

AI adoption evidence

What AI has actually done to productivity, by industry.

91 published results across 13 industries, each cited so you can check it. The studies that found no gain are here too, because a page that only carries supporting figures has told you nothing except which ones we chose.

None of these organisations are our clients. None of these numbers are ours. We publish no outcome figures of our own, and we would rather show you somebody else’s checkable result than our own unverifiable one.

The Australian position

Australia is further behind than the global coverage suggests.

The Australian Bureau of Statistics measured this directly, and the number is much lower than the survey figures quoted internationally. That gap is the single most useful fact on this page.

  • 12 per cent of Australian businesses used AI in 2024 to 2025

    Large businesses reached 35 per cent, up from 9 per cent in 2021 to 2022. Medium businesses reached 22 per cent. Small and micro businesses sat at 11 per cent. The gap between a large and a small business is now roughly threefold.

    Australian Bureau of Statistics, 2026

  • Adoption is concentrated in three industries

    Information, media and telecommunications led on 38 per cent. Professional, scientific and technical services and financial and insurance services both sat on 24 per cent. Financial and insurance services grew 24-fold from a 1 per cent base in three years.

    Australian Bureau of Statistics, 2026

  • Innovation activity predicts adoption better than size

    Among small businesses, those already innovation-active adopted AI at 19 per cent, almost five times the rate of small businesses undertaking no innovation activity at all.

    Australian Bureau of Statistics, 2026

Why the numbers disagree so much

Global surveys report 88 per cent of organisations using AI in some form. The ABS reports 12 per cent of Australian businesses. Both are correct. The global figure counts large enterprises answering a voluntary survey about any use at all, including one team trying one tool. The ABS surveyed nearly 7,000 businesses of every size through the Business Characteristics Survey. When somebody quotes an adoption rate at you, the sampling frame is doing most of the work.

The global picture

Nearly everyone has started. Very few can show a profit.

Read these four findings together. The distance between the first and the third is the entire commercial problem.

  • 88 per cent of surveyed organisations now use AI in some form

    Use is close to universal. Depth is not. Agent adoption remains in single digits across most business functions, and reported productivity gains are concentrated in a small leading cohort rather than spread across the surveyed population.

    Stanford HAI, AI Index Report, 2026

  • Measured gains vary by a factor of five depending on the task

    Roughly 14 to 15 per cent in customer support, about 26 per cent in software development, and far higher figures in marketing output. Gains shrink as the work requires more reasoning, judgement or accountability.

    Stanford HAI, AI Index Report, 2026

  • About 6 per cent of organisations can attribute real profit to AI

    Only that cohort attributes 5 per cent or more of earnings before interest and tax to AI use. Roughly 23 per cent are scaling agents somewhere in the business, but in any single function the ceiling sits at or below 10 per cent.

    McKinsey, The state of AI, 2026

  • Security and risk, not capability, is the reported blocker

    Nearly two-thirds of respondents named security and risk concerns as the leading barrier to scaling agentic AI, ahead of regulatory uncertainty and ahead of technical limitations.

    McKinsey, The state of AI, 2026

The evidence against

The findings that should make you slower.

These are not caveats we added for balance. Two are randomised controlled trials and one is a company reversing its own widely reported result.

  • 95 per cent of generative AI pilots showed no measurable return

    Drawn from 150 leader interviews, 350 employee surveys and 300 public implementations against an estimated 30 to 40 billion US dollars of enterprise spend. The authors describe the work as preliminary findings and it has not been peer reviewed, so treat the 95 per cent as a direction rather than a measurement. The direction is still the one worth planning around.

    MIT Project NANDA report, via Forbes, 2025

  • Experienced developers were 19 per cent slower with AI, and believed they were faster

    A randomised controlled trial of 16 experienced open-source maintainers across 246 real tasks in repositories they already knew well. They completed work 19 per cent slower with AI tools available, having forecast a 24 per cent speed-up, and still estimated afterwards that they had been 20 per cent faster. Self-reported productivity is not evidence.

    METR, 2025

  • The best-known customer service result was partly walked back

    Klarna's published figures, linked here, were real. The follow-through was harder: through 2025 the company rebalanced toward human agents after concluding that cost had been weighted too heavily against service quality. Those later statements were made to press rather than published by the company, so weigh them accordingly. The lesson is not that the automation failed, it is that the escalation path was underbuilt.

    Klarna, original press release, 2025

The finding we take most seriously

In the METR trial, experienced developers were measurably slower and still believed they had been faster. That result generalises well beyond software. It means a pilot assessed on how the participants felt about it will approve itself, whatever it did to the work. It is why we ask for a measurable baseline before an engagement starts, and why we treat a glowing user survey as the weakest evidence available rather than the strongest.

Sector by sector

13 industries, and what each one can actually prove.

Five of these are sectors we work in, marked below. The other eight are here because the evidence in them is instructive, not because we claim to serve them. The number beside each is how many cited results it carries.

Sector 1 of 13

Construction and property

What we build in construction

Where the gains are real

The measured gains are in reading, not building. Document-heavy preconstruction work (scope extraction, contract and specification review, progress reporting) is where the sector's evidence sits. Nothing credible shows AI improving on-site productivity at scale.

What the evidence does not support

This is the least-adopted sector in the evidence base. Barriers are not model quality: skills, systems integration and data quality account for the top three, which is a delivery problem rather than a technology one.

  • RICS member survey, 4,000 respondents worldwide

    Global

    Surveyed where construction and property professionals actually have AI in production, and what is stopping the rest.

    45 per cent reported no AI implementation and 34 per cent were in early pilots, leaving roughly 13 per cent using it regularly in specific processes and under 1 per cent embedded organisation-wide. Contract and document review was among the highest-value applications named, at 30 per cent.

    RICS, Artificial intelligence in construction, 2025

  • RICS, barriers to adoption

    Global

    Asked the same respondents what was actually blocking implementation.

    Lack of skilled personnel 46 per cent, integration with existing systems 37 per cent, data quality and availability 30 per cent, implementation cost 29 per cent, unclear return on investment 28 per cent.

    What it does not prove: Self-reported barriers from a professional-body survey. Useful for direction, not a measurement of any one firm.

    RICS, Artificial intelligence in construction, 2025

  • International Journal of Construction Management

    Peer-reviewed research

    Built and evaluated a generative AI model that assembles bid proposal components from a tender set.

    An average F1 score of 96.25 on extracting and structuring bid content, with expert evaluation finding reduced manual effort and better consistency between bids.

    What it does not prove: A research prototype measured on extraction accuracy, not a commercial deployment measured on won work or margin.

    Taylor and Francis, 2026

  • BAM Ireland

    Ireland

    Autodesk Construction IQ, a machine learning layer over BIM 360 issue and document data, ranking site quality and safety issues by predicted risk so supervisors work the highest-risk items first.

    BAM Ireland reported a 20% improvement in on-site quality and safety, and a 25% increase in the share of project staff time spent on high-risk issues, on projects generating 10,000 to 15,000 documents per building.

    What it does not prove: Figures are BAM Ireland's own, reported through its software vendor's programme, and are not independently audited. BAM's own account notes the first thing the system surfaced was inconsistent issue closure in BAM's own records rather than genuine site risk, so part of the gain is better data discipline rather than better prediction.

    AEC Magazine, 2019

  • Google DeepMind and Trane Technologies

    United States, two commercial facilities (not named in the paper)

    A reinforcement learning agent taking over control of live commercial cooling plant (chillers, pumps and towers) from the incumbent rule-based building control system, run as live experiments on two real facilities.

    Energy savings of approximately 9% and 13% at the two live experiment sites, measured against each site's existing controller.

    What it does not prove: Two sites only, and the paper is authored by the technology providers rather than an independent evaluator. Savings are relative to whatever control strategy each building already ran, so a well-tuned incumbent would leave less headroom. The facilities are not identified, which limits what a reader can check independently.

    arXiv preprint 2211.07357, 2022

  • Zillow Group

    United States

    Zillow Offers, an instant-buying business in which an automated valuation model forecast future selling prices and set the price Zillow paid for homes it bought, renovated and resold.

    Zillow announced on 2 November 2021 that it would wind the business down, after a $304 million inventory write-down in the third quarter, a further $240 million to $265 million expected in the fourth quarter, and a workforce reduction of approximately 25%. Chief executive Rich Barton said the company had determined "the unpredictability in forecasting home prices far exceeds what we anticipated".

    What it does not prove: The wind-down was driven by renovation and resale capacity constraints and an extreme housing market as well as by model error, so it is not purely a failure of the valuation model. It remains the clearest published case of a property business retiring a machine learning pricing model because the error was too expensive to carry on the balance sheet.

    Zillow Group third-quarter 2021 results release, 2021

  • Australian Bureau of Statistics

    Australia

    National business survey (Characteristics of Australian Business) measuring whether businesses used artificial intelligence in the workplace, published broken down by industry division.

    6% of Australian construction businesses reported using AI in 2024-25, against an all-business rate of 12% and 38% for information media and telecommunications. Construction ranked 14th of the 17 industry divisions published.

    What it does not prove: Adoption is not outcome. The survey records whether a business used AI at all, not whether it gained anything, and it captures general-purpose assistants alongside production systems. It is useful as a base rate for how early the sector actually is, not as evidence for or against value.

    Australian Bureau of Statistics, Characteristics of Australian Business, 2026

Back to all 13 sectors

Sector 2 of 13

Manufacturing and distribution

What we build in manufacturing

Where the gains are real

Two things are well evidenced: inline quality inspection using vision and sensor data, and predictive maintenance. Both work because the ground truth is unambiguous. A part either passes or it does not, and a machine either failed or it did not.

What the evidence does not support

Demand forecasting evidence is much weaker than vendors imply. Gains depend almost entirely on order-history quality and on whether planners trust the output enough to stop maintaining a parallel spreadsheet.

  • BMW Group

    Germany, and every BMW plant worldwide

    AIQX, an inline quality platform using conveyor-mounted cameras and sensors, with AI analysing the data in real time and pushing feedback to the operator on the line.

    Now standard across every BMW Group plant globally, covering variant identification, completeness checking and anomaly detection. The company is assessing making it available to suppliers.

    What it does not prove: BMW publishes the deployment scope but not a headline defect-rate figure. Worldwide rollout by a manufacturer of this size is the substantive signal here, not a percentage.

    BMW Group, 2023

  • United States Department of Energy, Federal Energy Management Program

    United States

    Established the baseline effect of moving from scheduled to predictive maintenance, which is the honest comparison for any AI maintenance business case.

    A properly functioning predictive maintenance programme returned 8 to 12 per cent savings over preventive maintenance alone, rising past 30 to 40 per cent where a facility had been relying on reactive maintenance.

    What it does not prove: Predates modern machine learning, and the two ranges are doing different work. The 8 to 12 per cent is the gain against an already competent programme. The 30 to 40 per cent is the gain against neglect. Which one a business is entitled to expect depends entirely on where it starts.

    US Department of Energy, O and M Best Practices Guide Release 3.0, 2010

  • Stanford HAI, AI Index

    Global

    Tracked where measured productivity gains actually landed across business functions.

    Gains fall as the work requires more reasoning. Structured, high-volume, verifiable tasks moved most, which is exactly the shape of inspection and maintenance work.

    Stanford HAI, AI Index Report, 2026

  • Toyota Motor Corporation

    Japan

    An in-house AI Platform that lets production line staff, rather than specialist AI engineers, build and deploy their own machine learning models against manufacturing line data.

    Nearly 1,200 employees use the platform across all 10 of Toyota's car and unit manufacturing factories, with more than 400 staff going through in-house training each year. Toyota reports over 10,000 man-hours saved per year, 10,000 models created in 2024 (up from 8,000 in 2023), and a 20% reduction in model creation time after a container image streaming change.

    What it does not prove: Written by Toyota's own AI group and published by its cloud vendor. The hours figure is an internal estimate and the post does not state how it was measured or against what baseline. Model count is an activity measure, not a value measure.

    Google Cloud blog, authored by Toyota Motor Corporation, 2024

  • Toyota Industries Corporation

    Japan, Nagakusa Plant

    Paint shop process data unified into a semantic layer on Azure with Sight Machine, analysing close to 400 process variables to find the drivers of paint defects on bumper production, replacing manual sight-based inspection judgement.

    A 25% reduction in seeding-related paint defects during the pilot, analysis cycles cut from 5 days to under 4 hours, and about 80% less preparation time for daily standup meetings.

    What it does not prove: A vendor-published customer story, not an independent evaluation. The 25% figure covers one defect type at one plant during a pilot phase, and the 18% carbon reduction quoted in the same story is an expectation rather than a measured result.

    Microsoft Customer Stories, 2026

  • Siemens

    China, Nanjing

    More than 50 artificial intelligence applications deployed alongside digital twins, modular automation and manufacturing operations management at Siemens' Nanjing Digital Industries factory.

    Measured against a 2022 baseline, Siemens reports lead times down 78%, time to market down 33%, productivity up 14% by 2024, field failures down 46%, and direct and energy-related carbon emissions down 28%. The site was named a World Economic Forum Global Lighthouse.

    What it does not prove: Siemens' own figures for its own plant. The AI applications went in alongside a wider rebuild of the factory, so the improvements cannot be attributed to AI on its own. The Lighthouse designation is a recognition programme, not an audit of the numbers.

    Siemens press release, 2026

  • Schneider Electric

    Indonesia, Batam

    Industrial IoT deployment at the Batam plant combining smart sensors, alarm prediction management, augmented reality and planning and scheduling tools, with the site used as a test bed for machine learning and predictive maintenance.

    A 44% reduction in machine downtime in one year and a 40% improvement in on-time delivery.

    What it does not prove: Company-published and not independently verified. Machine learning was one of several technologies introduced at the same time, so the downtime figure belongs to the whole programme rather than to AI. The release also describes the plant as a test bed for AI, which suggests the AI element was less mature than the sensing and scheduling work.

    Schneider Electric press release, 2019

  • McElheran, Li, Brynjolfsson, Kroff, Dinlersoz, Foster and Zolas, using the US Census Bureau Annual Business Survey

    United States, peer-reviewed working paper

    Analysis of AI use in production as reported by 850,000 US firms in the 2018 Annual Business Survey, the largest firm-level collection of its kind.

    Fewer than 6% of US firms used any of the measured AI technologies, though employment-weighted adoption was just over 18%. Manufacturing was one of the two leading sectors at roughly 12% adoption, while construction and retail trade lagged at roughly 4%.

    What it does not prove: The survey year is 2018, before generative AI, so this describes the starting point rather than today's adoption. It measures use, not benefit: the paper does not establish that adopting firms performed better.

    NBER Working Paper 31788, 2023

Back to all 13 sectors

Sector 3 of 13

Logistics and warehousing

What we build in logistics

Where the gains are real

The strongest published numbers in any sector, because logistics measures everything already. Where a business can already count items per hour and dwell time to the minute, an improvement is provable rather than argued.

What the evidence does not support

Almost all headline figures come from operators with capital budgets in the hundreds of millions. The document-handling and exception-detection work transfers to a mid-size operator. The robotics numbers do not.

  • Amazon

    United States, global network

    Sequoia, a containerised inventory system combining robotics with computer vision for identification and storage.

    Up to 25 per cent faster order processing through a fulfilment centre, and inventory identified and stored up to 75 per cent faster.

    What it does not prove: Company-published figures, not independently audited.

    Amazon, 2023

  • Stanford HAI, AI Index

    Global

    Measured productivity effects in customer support work, which is the shape of a freight service desk answering where a consignment is and why it is late.

    Roughly 14 to 15 per cent, with the largest gains going to less experienced staff and the smallest to the most experienced.

    What it does not prove: Not measured in logistics specifically. Included because this, rather than the robotics figure above, is the number a mid-size operator should build a business case on.

    Stanford HAI, AI Index Report, 2026

  • Rio Tinto

    Western Australia, Pilbara

    AutoHaul, autonomous heavy-haul freight trains carrying iron ore from 16 mines to four port terminals, supervised remotely from an operations centre in Perth.

    More than 1 million kilometres travelled autonomously as at December 2018, across a network of about 200 locomotives and more than 1,700km of track, with an average return journey of about 800km and an average cycle time of about 40 hours including loading and dumping. Rio Tinto puts the programme cost at $940 million.

    What it does not prove: Rio Tinto's own release, and it claims potential rather than realised productivity gains: it says the deployment shows significant potential to improve productivity and reduce bottlenecks, without publishing a cycle-time or cost-per-tonne improvement. This is sensing and control automation rather than a learned model, and Rio Tinto stated it expected no redundancies in 2019 from the deployment.

    Rio Tinto media release, 2018

  • Union Pacific Railroad

    United States

    Machine Vision Systems photographing rolling stock, with the images processed by machine learning and AI to identify abnormal or defective components, feeding a network of more than 7,000 wayside detection devices.

    The wayside detector network generates more than 16 million data points every day. Union Pacific reports total derailments down 21% comparing 2022 with 2019, and track-caused derailments down 55% over the previous 10 years.

    What it does not prove: Union Pacific's own reporting. The derailment reductions cover the whole safety and inspection programme, including track maintenance practice and physical detectors, so they cannot be attributed to machine vision alone. No before-and-after figure is published for the machine vision component on its own.

    Union Pacific, 2023

  • Ocado Group

    United Kingdom

    On Grid Robotic Pick, robotic arms picking items directly from the storage grid in Ocado Smart Platform customer fulfilment centres, rolled out alongside Automated Frameload.

    Overall customer fulfilment centre productivity across the Ocado Smart Platform rose 9.1 per cent in the 2024 financial year, with more than 30 per cent of Luton volumes picked robotically by year end and rollout contracts signed with the majority of partners.

    What it does not prove: Company-reported, and Ocado attributes the productivity gain to higher volume utilisation as well as to robotic picking, so the robotics contribution cannot be isolated from the figure. Luton is the most advanced site rather than a typical one.

    Ocado Group, full year results 2024, 2025

  • Australian Bureau of Statistics

    Australia

    National business survey (Characteristics of Australian Business) measuring whether businesses used artificial intelligence in the workplace, published broken down by industry division.

    1% of Australian transport, postal and warehousing businesses reported using AI in 2024-25, the lowest of the 17 industry divisions published, against an all-business rate of 12%.

    What it does not prove: Adoption is not outcome, and the figure covers all transport, postal and warehousing businesses including very small operators, so it is dominated by owner-drivers and small fleets rather than by the large logistics operators most likely to have deployed something. Read it as a base rate for the sector, not as evidence that AI does not work in logistics.

    Australian Bureau of Statistics, Characteristics of Australian Business, 2026

Back to all 13 sectors

Sector 4 of 13

Professional services

What we build in professional services

Where the gains are real

This is the sector with genuine independent benchmarking, and the results are stronger than most partners expect. On defined document tasks, current tools now beat a lawyer baseline. On tasks needing judgement across a whole matter, they do not.

What the evidence does not support

The benchmark that produced these numbers was declined by three of the largest legal research vendors, so it measures the tools that opted in. Separately, none of this evidence addresses privilege, conflicts or confidentiality, which is where firm-level deployments actually stall.

  • Vals Legal AI Report, with a human lawyer control group

    Global, using real law firm data

    The first independent benchmark of legal AI tools across seven tasks lawyers actually do, including data extraction, document question answering, summarisation, redlining and chronology generation.

    AI outperformed the lawyer baseline in four of the seven areas tested. On document question answering the best tool scored 94.8 per cent against a lawyer baseline of 70.1 per cent.

    What it does not prove: Thomson Reuters, LexisNexis and vLex did not participate. The lawyer baseline is a control group, not a firm's best available specialist.

    Vals AI, via LawSites, 2025

  • Vals Legal AI Report, legal research extension

    Global

    Extended the benchmark to legal research specifically, comparing three legal AI systems and one general foundation model against a lawyer baseline.

    All AI systems landed around 80 per cent accuracy against a lawyer baseline of 71 per cent.

    What it does not prove: 80 per cent accuracy means one answer in five is wrong. That is a review workflow, not an answer service.

    Vals AI, via LawSites, 2025

  • Noy and Zhang, published in Science

    Peer-reviewed research

    A randomised experiment on mid-level professional writing tasks of the kind consultants, marketers and analysts do daily.

    Time taken fell about 40 per cent and assessed output quality rose about 18 per cent, with the largest gains going to the lower-performing participants.

    What it does not prove: Measured on discrete 20 to 30 minute tasks. A task speed-up is not the same as a billable-hour or realisation improvement.

    Science, 2023

  • Brynjolfsson, Li and Raymond study of a Fortune 500 software firm's customer support operation

    United States (peer-reviewed working paper)

    Staggered rollout of a generative AI conversational assistant to customer support agents, with agents who had not yet received it acting as the comparison group.

    Across 5,179 customer support agents, access to the assistant raised productivity, measured as issues resolved per hour, by 14 per cent on average. Novice and low-skilled workers improved by 34 per cent, while experienced and highly skilled workers saw minimal effect.

    What it does not prove: One firm, one job type, and an early-generation assistant, so the size of the gain is specific to that setting. The paper was revised before journal publication in the Quarterly Journal of Economics, so figures differ slightly between versions and you should quote the version you link to.

    National Bureau of Economic Research, 2023

  • Ashurst (now Ashurst Perkins Coie)

    Global law firm, 23 offices across 14 countries

    Firm wide trials of three generative AI tools across legal and business services tasks, including a blind study in which an expert panel scored AI drafts against lawyer drafts without knowing which was which.

    411 partners, lawyers and staff took part between November 2023 and March 2024. Reported time savings of roughly 80 per cent on drafting UK corporate filings, 59 per cent on industry research reports and 45 per cent on first drafts of legal briefings. In the blind study, AI outputs averaged 3.0 out of 5 for accuracy against 3.5 for lawyer written outputs, and the panel correctly identified every human written output as human.

    What it does not prove: Published by the firm itself and not independently audited. The savings are task level, measured on set exercises rather than on billed client matters, and the same report shows AI accuracy scoring below the firm's own lawyers.

    Ashurst, 2024

  • METR (Model Evaluation and Threat Research)

    Randomised controlled trial on mature open source repositories

    Experienced open source maintainers were randomly permitted or not permitted to use AI tools on real issues in repositories they already knew well, with completion times recorded.

    Across 16 developers and 246 issues, developers took 19 per cent longer to complete work when AI tools were allowed. The same developers had expected AI to speed them up by 24 per cent beforehand, and still believed they had been 20 per cent faster afterwards.

    What it does not prove: Small sample in a narrow setting: expert maintainers on codebases they know deeply. METR explicitly states the result is not evidence that AI fails to speed up most developers, and that it reflects the tools available in early 2025. The value here is the gap between perceived and measured speed, not the 19 per cent itself.

    METR, 2025

  • Linklaters

    United Kingdom, English law

    The LinksAI English law benchmark: 50 questions across 10 practice areas, marked out of 10 (5 for substance, 3 for citations, 2 for clarity) against the standard of a competent mid-level lawyer with two years post qualification experience.

    In the February 2025 round, OpenAI o1 scored 6.4 out of 10 and Gemini 2.0 scored 6.0, against a best score of 4.4 out of 10 in the October 2023 round. Linklaters concludes the models should not be used for English law advice without expert human supervision, though they may be useful for a first draft or a cross-check in well-known areas of law.

    What it does not prove: A 50 question benchmark designed and marked by one firm, not a measure of work delivered to clients. Model versions change quickly, so the scores date fast.

    Linklaters, 2025

  • Deloitte Australia and the Department of Employment and Workplace Relations

    Australia, Commonwealth government

    An independent assurance review of the Targeted Compliance Framework, in which a generative AI tool chain was used to assess whether system code could be traced to business requirements.

    The report published on 3 October 2025 was withdrawn after fabricated academic references and an invented quotation attributed to a Federal Court judgment were identified. The revised report discloses the use of "a generative AI large language model (Azure OpenAI GPT 4o) based tool chain" licensed by the department. Deloitte agreed to repay the final instalment of a contract worth just under A$440,000. The department published a further corrected version on 3 February 2026.

    What it does not prove: The department's page records only that corrections were made and that the report replaces the version published on 3 October 2025. The detail of the fabricated citations and the repayment comes from reporting and from Deloitte's own statements, not from a published audit, and the refund amount was not disclosed at the time.

    Department of Employment and Workplace Relations, 2025

Back to all 13 sectors

Sector 5 of 13

Healthcare and community services

What we build in healthcare and community

Where the gains are real

The clearest evidence in any sector, and it is entirely administrative. Ambient documentation gives clinicians time back without touching a clinical decision, which is precisely why it cleared governance and scaled where other health AI did not.

What the evidence does not support

None of this evidence is clinical, and we do not extend it there. Published results also depend on quality assurance running continuously over generated notes, which is an operating cost rather than a one-off implementation.

  • The Permanente Medical Group and Kaiser Permanente

    United States, Northern California

    Ambient AI scribes generating draft clinical documentation from the consultation itself, for physician review and editing.

    Across 7,260 physicians and roughly 2.5 million patient encounters between October 2023 and December 2024, an estimated 15,791 hours of documentation time saved, equivalent to about 1,794 eight-hour working days.

    What it does not prove: Documentation time only. The publication also describes the ongoing quality assurance the organisation had to run alongside it.

    NEJM Catalyst, 2025

  • Stanford HAI, AI Index

    Global

    Tracked clinician-reported effects of automatic clinical note generation as adoption scaled through 2025.

    Physicians reported up to 83 per cent less time spent writing notes, alongside meaningful reductions in reported burnout.

    What it does not prove: Clinician self-report, which tends to overstate against measured time.

    Stanford HAI, AI Index Report, 2026

  • Sutter Health

    United States, Northern and Central California

    Evaluation of an ambient AI documentation platform, using electronic health record audit metrics plus pre and post implementation surveys.

    Among 100 clinicians (92 with EHR metrics, 57 completing both surveys), mean time in notes per appointment fell from 6.2 to 5.3 minutes in the 3 months after the April 2024 launch. After hours EHR time did not change significantly, moving from 38.2 to 39.8 minutes. Clinicians reporting they could give patients undivided attention rose from 57.9 per cent to 93.0 per cent.

    What it does not prove: Single organisation, purposive sampling of volunteer clinicians, and only 3 months of follow up. The null result on after hours work matters as much as the note time saving: the documentation burden moved, it did not leave the evening.

    JAMA Network Open, 2025

  • Mass General Brigham and Emory Healthcare

    United States, Massachusetts and Georgia

    Survey study of pilot users of ambient documentation technology, measuring self-reported burnout and documentation related wellbeing.

    Across more than 1,400 physicians and advanced practice providers (873 at Mass General Brigham, 557 at Emory Healthcare), burnout prevalence fell by 21.2 percentage points at 84 days at Mass General Brigham, and documentation related wellbeing rose by 30.7 percentage points at 60 days at Emory.

    What it does not prove: Self-reported survey measures, not measured time or clinical output, and published by the health system about its own study. Response rates were roughly 30 per cent at 42 days and 22 per cent at 84 days at Mass General Brigham and 11 per cent at Emory. The authors state the findings likely represent the experience of more enthusiastic pilot users.

    Mass General Brigham (reporting its JAMA Network Open study), 2025

  • Mass General Brigham

    United States, 5 hospitals

    Comparison of clinicians using AI scribes against a control group of clinicians who did not, using electronic health record audit log data over more than 2 years.

    Across more than 1,800 AI scribe users and 6,770 control clinicians, EHR use fell by 13 minutes a day (a 3 per cent relative decrease) and documentation time by 16 minutes a day (a 10 per cent relative decrease), alongside 0.5 additional patient visits a week and about US$167 a month in additional revenue per clinician. Only 32 per cent of users reached the adoption level (more than half of visits) associated with the larger reductions.

    What it does not prove: Observational rather than randomised, and summarised here by the health system that ran it. The authors note the time reductions are too modest to fully account for the improvements in burnout reported in other studies, which is the point worth carrying to a business case.

    Mass General Brigham (reporting its JAMA study), 2026

  • Gold Coast Hospital and Health Service

    Queensland, Australia

    A 16 week outpatient trial of an AI-enabled ambient listening scribe across medical officers, registrars and clinical nurses, with blinded scoring of note quality and staff and patient surveys.

    Across 100 clinicians and 7,499 consultations between July and December 2024, 58 per cent of AI generated notes were accepted verbatim. Blinded quality scoring put AI notes at 37.06 out of 40 against 34.6 for clinician written notes. 84 per cent of staff reported improved efficiency and 68 per cent of patients said clinicians spent more time in direct conversation. However 47 per cent of staff reported hallucinations, 20 per cent said they occurred frequently, and 16 per cent observed bias in outputs.

    What it does not prove: Single site, and the note quality comparison rests on only 18 note pairs. Staff survey response rate was 43 per cent with just 22 patient surveys, the efficiency claims are self-reported rather than timed, and no economic evaluation was done.

    BMC Health Services Research, 2026

  • Michigan Medicine, external validation of the Epic Sepsis Model

    United States, Michigan

    Independent external validation of a proprietary sepsis prediction model that had been deployed across hundreds of United States hospitals, run by researchers with no commercial interest in the model.

    Across 38,455 hospitalisations among 27,697 patients between 6 December 2018 and 20 October 2019, the model achieved an area under the curve of 0.63 (95 per cent CI 0.62 to 0.64). It failed to identify 1,709 of 2,552 patients with sepsis (67 per cent), while generating alerts on 6,971 of all 38,455 hospitalised patients (18 per cent).

    What it does not prove: A predictive model rather than a generative AI system, and a single academic health system. It is included because it is the clearest published example of a widely deployed clinical model performing far below its marketed claims once someone outside the vendor measured it.

    JAMA Internal Medicine, 2021

Back to all 13 sectors

Sector 6 of 13

Retail and consumer goods

Not a sector we position in. Published here for the pattern.

Where the gains are real

Retail's returns are showing up in advertising yield and merchandising rather than in the shopping assistants that get the attention. The money is in deciding what to stock, price and promote.

What the evidence does not support

Customer-facing assistants report large usage numbers and very few conversion or margin figures. Usage is not value.

  • Zalando

    Europe

    AI applied across advertising optimisation, content production, logistics accuracy and a customer-facing shopping assistant.

    Retail media revenue grew 42 per cent in 2025, attributed in part to AI-optimised advertising performance, against group revenue growth of about 17 per cent to 12.3 billion euros. Close to 10 million customers used the assistant for advice.

    What it does not prove: Company-reported and multi-causal. Retail media grows for reasons beyond AI, and Zalando does not isolate the AI contribution.

    Zalando investor relations, 2026

  • Walmart

    United States

    Rolled AI assistant tooling into the associate app and stood up agents across operational functions, particularly inventory management.

    Reached about 1.5 million associates, with the company crediting automation in inventory management for supporting sales growth through a tariff-affected year.

    What it does not prove: Deployment scale and management attribution, not a measured productivity figure.

    CIO Dive, reporting company earnings commentary, 2026

  • Ocado Group

    United Kingdom

    Deep learning demand forecasting inside the Ocado Smart Platform, driving automatic stock ordering for online grocery.

    Ocado states its platform generates more than 70 million supply chain forecasts each day, that its deep learning forecasts are "up to 40% more accurate than traditional retailer systems, designed for bricks and mortar operations", and that "up to 98% of stock [is] ordered automatically with minimal manual intervention".

    What it does not prove: Published by the company, not independently audited. The 40 per cent figure is an "up to" ceiling rather than an average, and the comparison baseline (forecasting systems built for physical stores) is chosen by Ocado.

    Ocado Group newsroom, 2025

  • Amazon

    United States and global Amazon stores

    Rufus, a generative AI shopping assistant built into the Amazon store and app, answering product questions and now able to complete purchases.

    In its fourth quarter 2025 results release Amazon stated that Rufus "was used by 300 million+ customers and saw an even stronger response than anticipated, helping deliver nearly $12 billion in incremental annualized sales last year".

    What it does not prove: A company figure in an earnings release. Amazon does not publish how it attributes incremental sales to Rufus, and no counterfactual or control group is disclosed, so the 12 billion dollars is an internal attribution rather than a measured lift.

    Amazon, fourth quarter 2025 results, 2026

  • Ingka Group (IKEA)

    Global IKEA retail operations

    Billie, an AI customer service assistant handling routine enquiries, with the call centre workforce retrained as remote interior design advisers rather than made redundant.

    Ingka reports Billie "resolved approximately 47% of customer enquiries", covering 3.2 million interactions and close to EUR 13 million in savings, while 8,500 call centre co-workers gained new competencies. Remote customer meeting points reached EUR 1.3 billion of sales, or 3.3 per cent of total sales, at the end of FY22.

    What it does not prove: Company-published and now several years old. The EUR 1.3 billion is the whole remote selling channel, not a benefit attributable to the chatbot; the two numbers are reported alongside each other rather than causally linked. Ingka has since made unrelated redundancies, so this is not evidence that reskilling permanently avoided job losses.

    Ingka Group newsroom, 2023

  • Bunnings Group

    Australia, Victoria and New South Wales

    Facial recognition applied to CCTV, capturing the face of every person entering a store and matching it against a database of individuals linked to violent incidents and theft.

    On 19 November 2024 the Privacy Commissioner determined that Bunnings breached the Privacy Act by collecting sensitive biometric information without consent across 63 stores between November 2018 and November 2021, affecting likely hundreds of thousands of people, and ordered the practice not to be repeated. The Commissioner's finding was that facial recognition was "the most intrusive option, disproportionately interfering with the privacy of everyone".

    What it does not prove: This is a compliance finding, not a performance measure: Bunnings never published effectiveness figures for the system. It was also partly overturned. In a statement published 5 March 2026 the Privacy Commissioner records that the Administrative Review Tribunal's Guidance and Appeals Panel found Bunnings was entitled to use the technology for that purpose, while leaving the notification and governance findings undisturbed. Read it as evidence about consent and notification obligations, not as a verdict on whether the technology works.

    Office of the Australian Information Commissioner, 2024

  • McDonald's and IBM

    United States

    Automated order taking, a voice AI system for drive-thru ordering, tested with IBM over roughly two years.

    McDonald's ended the IBM automated order taking partnership and switched the technology off in all test restaurants no later than 26 July 2024. The company said the test "has given us the confidence that a voice ordering solution for drive-thru will be part of our restaurants' future" while it explored options more broadly.

    What it does not prove: Trade press reporting on an internal franchisee memo, not a McDonald's publication. No accuracy, throughput or labour figures were ever released for the pilot, so the reasons for ending it are not documented in measured terms. McDonald's has continued to pursue voice ordering with other partners.

    Restaurant Dive, 2024

Back to all 13 sectors

Sector 7 of 13

Financial services and insurance

Not a sector we position in. Published here for the pattern.

Where the gains are real

The most-cited productivity numbers in any sector come from customer service in this one. They are real, and the sector is also where the clearest published correction happened, which makes it the most instructive evidence available.

What the evidence does not support

This is the cautionary sector. The headline result was followed by a partial reversal, and the reason was service quality, not model capability.

  • Klarna

    Global

    An AI assistant handling front-line customer service enquiries.

    In its first month it handled two-thirds of service chats across 2.3 million conversations, cut average resolution time from 11 minutes to under two minutes, reduced repeat enquiries by 25 per cent, and held customer satisfaction level with human agents. The company estimated a 40 million US dollar profit improvement.

    What it does not prove: The 700-agent equivalence is a workload calculation, not a redundancy count, and the 40 million dollar figure was a forward projection rather than an audited saving. Read this alongside the reversal noted in the evidence against.

    Klarna, 2024

  • Morgan Stanley

    United States

    Gave financial advisers retrieval over the firm's research and document library, plus automatic generation of client meeting notes.

    Regular use reached 98 per cent of adviser teams, across a population of roughly 16,000 advisers.

    What it does not prove: Adoption, not productivity. Included because near-universal voluntary use in a compliance-heavy population is itself hard to achieve.

    Morgan Stanley, 2024

  • Australian Bureau of Statistics

    Australia

    Measured AI adoption by industry across Australian businesses.

    Financial and insurance services reached 24 per cent adoption, a 24-fold increase from a 1 per cent base in 2021 to 2022. The fastest-moving sector in the Australian economy from the lowest start.

    Australian Bureau of Statistics, 2026

  • Commonwealth Bank of Australia

    Australia

    Generative AI powered fraud and scam alerts, AI-assisted app messaging in place of phone queues, and AI-supported annual credit reviews.

    CBA reported "a 50 per cent reduction in customer scam losses", "a 30 per cent drop in customer-reported frauds" from generative AI powered alerts, call centre wait times down 40 per cent over the prior financial year, and annual credit reviews cut from roughly 14 hours to two hours. It was sending 20,000 proactive warning alerts a day, scaling towards 35,000.

    What it does not prove: All figures are the bank's own, published in a media release rather than audited. Several controls changed in the same period, so attributing the falls in scam losses and reported fraud to AI is CBA's judgement, not a controlled result.

    CommBank Newsroom, 2024

  • Commonwealth Bank of Australia and the Finance Sector Union

    Australia

    An AI voice bot in CBA's Customer Service Direct business, whose reported effect on call volumes was used to justify making 45 customer service roles redundant.

    After the union raised a dispute at the Fair Work Commission, CBA reversed all 45 redundancies and accepted the roles were not required to go. The union records that management "admitted they didn't properly consider that an increase in calls would continue over a number of months", against the bank's earlier position that the voice bot had reduced call volumes.

    What it does not prove: Published by the union that ran the dispute, so it is one party's account of a contested matter. The reversal and CBA's apology were widely reported at the time, but the specific claim about the size of the call volume reduction comes from the union, not from CBA. Treat this as evidence about how easily a deployment's benefit can be overstated internally, not as a measurement of the voice bot itself.

    Finance Sector Union, 2025

  • National Australia Bank

    Australia

    "Customer Brain", a decisioning engine that selects next best actions and proactive prompts across NAB's customer channels.

    NAB reports the Brain "helps guide more than 50 million interactions" a month, has "helped drive a 40% uplift in customer engagement" in just over a year, and that proactive prompts about term deposit expiries "resulted in approximately $92 million in retained deposits". It runs over 2,000 adaptive models powering 220 next best actions.

    What it does not prove: Bank-published and not externally verified. "Uplift in customer engagement" is an internal marketing metric with no stated baseline, and the 92 million dollars of retained deposits is NAB's own attribution rather than a controlled comparison against customers who received no prompt.

    NAB News, 2025

  • Bank of America

    United States

    Erica, a virtual assistant for retail customers, plus "Erica for Employees" used internally for IT and workplace support.

    Bank of America reported Erica had passed 3 billion client interactions and nearly 50 million users since its 2018 launch, averaging more than 58 million interactions a month and 1.7 billion proactive personalised insights delivered. Internally, 90 per cent of employees use Erica for Employees, which the bank credits with a 50 per cent reduction in IT service desk calls.

    What it does not prove: A company press release. Almost every figure is a volume measure, which shows adoption rather than benefit. The 50 per cent fall in IT service desk calls is the only outcome number and no baseline period, sample or method is published alongside it.

    Bank of America Newsroom, 2025

  • Bank of England and Financial Conduct Authority

    United Kingdom

    The regulators' third joint survey of AI and machine learning use across UK financial services firms.

    The 2024 survey found 75 per cent of firms are already using AI, with a further 10 per cent planning to within three years, against 58 per cent and 14 per cent in the 2022 survey. Foundation models, including large language models, account for 17 per cent of all AI use cases reported.

    What it does not prove: Adoption self-reported by firms to their regulator, so it measures declared use rather than realised benefit, and says nothing about whether any deployment paid for itself. The sample is UK-regulated firms only.

    Financial Conduct Authority research note, 2024

Back to all 13 sectors

Sector 8 of 13

Energy and utilities

Not a sector we position in. Published here for the pattern.

Where the gains are real

Forecasting is where the value is. Knowing what the grid or the wind farm will do in 36 hours changes what can be committed and sold, and the commercial gain comes from the commitment rather than from generating more energy.

What the evidence does not support

Grid figures are mostly published by the utility or its technology partner. Independent replication is rare, and results depend heavily on the sensor estate already being in place.

  • Google DeepMind and Google

    Central United States wind farms

    Machine learning forecasting wind power output 36 hours ahead, allowing generation to be committed into the market in advance.

    Roughly a 20 per cent increase in the value of the wind energy produced, relative to a baseline with no time-based commitments.

    What it does not prove: Value, not output. The turbines did not generate more, the energy was worth more because it could be sold ahead.

    Google DeepMind, 2019

  • Enel and E.ON, reported deployments

    Europe

    Sensor telemetry with machine learning applied to cable and network asset monitoring, and predictive rather than scheduled maintenance.

    Reported outage reductions of about 15 per cent on monitored cables, with predictive maintenance approaches assessed as capable of reducing grid outages by up to 30 per cent against scheduled maintenance.

    What it does not prove: Secondary reporting of utility deployments, and the 30 per cent figure is an assessed potential rather than a measured result.

    EY, power and utilities analysis, 2025

  • Google and DeepMind

    United States and other Google data centre sites

    A reinforcement learning control system given direct authority over the cooling plant, working inside a human-defined safety envelope, replacing operator-set cooling schedules.

    Around 30 per cent average energy savings on cooling, measured as energy input per unit of cooling delivered, across multiple data centres after roughly nine months of live operation (improving from 12 per cent in September 2017 to about 30 per cent by July 2018). The earlier recommendation-only version reported a 40 per cent reduction in cooling energy and a 15 per cent reduction in overall PUE overhead.

    What it does not prove: Published by Google about its own facilities, with no independent audit and no release of the underlying data. The baseline is Google's already heavily optimised cooling operation, so the percentage should not be read across to a typical commercial data centre.

    Google DeepMind, 2018

  • National Energy System Operator (formerly National Grid ESO) and Open Climate Fix

    Great Britain

    A machine learning solar generation nowcast built from satellite imagery, numerical weather prediction and live output from thousands of PV systems, feeding the control room's dynamic reserve setting.

    Open Climate Fix reports the service halved NESO's solar forecast error and cut total system demand forecast error by 8 per cent, which it associates with more than 30 million pounds and 300,000 tonnes of CO2 avoided per year through lower reserve holding. An earlier phase of the same project reported a mean absolute error of 233 MW against 650 MW for the incumbent forecast.

    What it does not prove: Published by the forecast supplier rather than by NESO, and the page does not state the measurement period or the method behind the financial and CO2 figures. NESO's own pages describing the project returned HTTP 403 and could not be checked.

    Open Climate Fix, 2025

  • Australian Renewable Energy Agency, the Australian Energy Market Operator and 11 wind and solar generators

    Australia, National Electricity Market

    A 9.4 million dollar trial in which wind and solar farms used weather data, site conditions and machine learning to submit their own five minute ahead output forecasts into AEMO's central dispatch engine, rather than relying on AEMO's AWEFS and ASEFS forecasting systems.

    On mean absolute error, participant self-forecasts beat AEMO's wind forecasts by between 11.7 and 19.8 per cent, and beat AEMO's solar forecasts by between 13.1 and 18.2 per cent. Average normalised mean absolute error was 2.6 per cent for wind farms and 3.0 per cent for solar farms over six months of data. Some participants reported that about 18 per cent improvement on MAE was the point at which further accuracy stopped paying for itself.

    What it does not prove: The head-to-head accuracy comparison came from a limited number of participants who supplied their own error statistics. Gains were uneven: on RMSE, solar improvement ranged from 2.8 to 21.2 per cent. The financial benefit runs through the causer pays FCAS cost recovery mechanism and was modelled rather than directly measured.

    Australian Renewable Energy Agency, report prepared by GHD Advisory, 2021

  • Alec Brandon, Christopher Clapp, John List, Robert Metcalfe and Michael Price, using data supplied by Opower

    United States

    Two randomised field experiments installing learning smart thermostats in volunteer households, testing whether the energy savings manufacturers advertise show up in actual metered consumption.

    Across 1,379 households in the electricity sample and 1,369 in the natural gas sample, using 18 months of data and more than 16 million hourly electricity and daily gas observations, the point estimates were a 0.09 per cent decrease in electricity use and a 1.70 per cent increase in gas use, neither statistically nor economically significant. Manufacturer claims at the time ranged from 10 to 23 per cent. Analysis of nearly four million thermostat system events attributes the gap to household override behaviour that engineering models assume away.

    What it does not prove: An NBER working paper (September 2022, revised November 2024), which is not peer reviewed. Households volunteered to take part, so they are not a random sample of the population. Opower supplied the data under a non-disclosure agreement with a right of factual review, though the authors state the agreement preserved their independence.

    National Bureau of Economic Research working paper 30482, 2022

Back to all 13 sectors

Sector 9 of 13

Mining and resources

Not a sector we position in. Published here for the pattern.

Where the gains are real

Australia holds the longest-running production evidence in the world here, and it is measured in unit cost and machine utilisation rather than in software metrics. Autonomy pays because a machine that does not need a shift change runs more hours.

What the evidence does not support

These are 15-year capital programmes on fixed haul routes in a controlled environment. The unit economics do not transfer to a variable or public setting.

  • Rio Tinto

    Pilbara, Western Australia

    Autonomous haulage across an iron ore truck fleet, supervised centrally rather than driven.

    About 15 per cent lower load and haul unit costs than the conventional equivalent, with each autonomous truck operating roughly 700 hours more per year.

    What it does not prove: Operator-reported against its own conventional baseline. The comparison is internal, not independent.

    Rio Tinto, via Global Mining Guidelines Group, 2021

  • Rio Tinto AutoHaul

    Pilbara, Western Australia

    Autonomous heavy-haul rail moving iron ore to port without drivers on board.

    More than 7 million kilometres run since 2019, supporting 326.2 million tonnes railed in the Pilbara in 2025, and removing close to 1.5 million kilometres of road travel a year previously needed to move drivers to and from trains.

    What it does not prove: The safety benefit of removing that road travel is more clearly established than any productivity figure.

    International Railway Journal, 2025

  • BHP

    Chile, Escondida copper operation

    AI and machine learning applied across the Escondida processing plants, using real-time plant data to predict conditions and recommend operator adjustments to concentrator and desalination water and energy use.

    BHP states in its FY2024 Annual Report that the application of AI at the Escondida processing plants has helped save more than three gigalitres of water and 118 gigawatt hours of energy since FY2022.

    What it does not prove: A single narrative line in a company annual report. No methodology is given for separating the AI contribution from other plant and process changes across the same two years, and no copper recovery or throughput figure is attached to it. BHP's own web pages carrying the same claim were unreachable, so the filing is the checkable source.

    BHP Annual Report 2024, filed with the US Securities and Exchange Commission, 2024

  • Vale

    Brazil, Itabira, Minas Gerais

    Rebuild of the Conceicao 2 iron ore beneficiation plant around automated instrumentation and an AI supervisory layer that monitors and adjusts more than 400 process variables in real time according to ore and product characteristics.

    Vale reports productivity up 25 per cent, direct reduction pellet feed output up 40 per cent, iron content in tailings down 26 per cent in 2026, and 92 per cent of process water recirculated. Plant capacity moved from 9 million tonnes produced in 2024 to a rated 11.2 million tonnes a year. The work took 1.5 years, involved 51 separate changes, added roughly 7,300 automated instruments and more than 100 cameras, and retrained 122 operators and technicians.

    What it does not prove: Company-published and not independently audited. The plant was rebuilt as a whole through 51 separate interventions, so the gains cannot be attributed to the AI layer on its own. The tailings and productivity figures are Vale's own measurements against its own prior baseline.

    Vale, 2026

  • Fortescue

    Australia, Pilbara, Western Australia

    Autonomous haulage across three iron ore mines, with trucks dispatched and routed by a central system and supervised from a remote operations centre in Perth.

    By July 2021 the autonomous fleet had moved two billion tonnes of material, double the volume reported at the one billion tonne milestone in September 2019, using 193 trucks (79 at Solomon, 74 at Christmas Creek and 40 at Cloudbreak) and travelling more than 70 million kilometres.

    What it does not prove: A company announcement reporting volume and distance milestones, not a measured productivity or safety result. No comparison against manned haulage is given and no independent verification is offered. Productivity claims of around 30 per cent that circulate alongside this milestone appear in trade press, not in this release.

    Fortescue, 2021

  • University of Queensland Sustainable Minerals Institute and University of Pittsburgh, for ACARP

    Australia, drawing on Western Australian regulator incident data

    An independent review of the measured benefits and the credible failure modes of mining equipment automation, using operator incident data and 53 incident summaries published by the Western Australian Department of Mines, Industry Regulation and Safety between January 2010 and May 2021.

    The review reports that BHP's Jimblebar mine recorded a fall of more than 90 per cent in the overall haul truck incident rate across the four years spanning the introduction of autonomous haulage, and that Rio Tinto reported an order of magnitude difference in collision near misses between autonomous and manual sites. Against that, it identifies recurring failure modes in the regulator's own records: 18 incidents where a manually operated vehicle encroached into an autonomous truck's permission line, seven of which ended in collision, plus repeated loss-of-traction events on wet or overwatered roads and one communications failure in which a truck reversed into a parked truck.

    What it does not prove: The Jimblebar and Rio Tinto safety figures are company-reported numbers cited by the review, not independently reproduced by the authors. The incident analysis, which is the authors' own work, is the more robust half. The review is funded by the Australian coal industry research programme ACARP.

    University of Queensland Sustainable Minerals Institute, ACARP project C34026 white paper, 2024

Back to all 13 sectors

Sector 10 of 13

Agriculture and food

Not a sector we position in. Published here for the pattern.

Where the gains are real

The best-evidenced input reduction anywhere. Computer vision deciding what to spray, plant by plant, at working speed, is now running across millions of hectares with published season-level results.

What the evidence does not support

Results depend on weed pressure, crop and season, and vary widely year to year. The equipment cost puts it out of reach of smaller operations without contracting arrangements.

  • John Deere, See and Spray

    United States, farmer fleets

    Boom-mounted cameras with machine learning identifying weeds at up to 15 miles per hour and triggering individual nozzles, rather than spraying the whole field.

    Across 5 million acres in the 2025 season, non-residual herbicide use fell by an average of close to 50 per cent, saving an estimated 31 million gallons of herbicide mix. The 2024 season averaged 59 per cent savings across a smaller area.

    What it does not prove: Manufacturer-published from customer fleet data. The variance between the 2024 and 2025 averages shows how much the season drives the number.

    John Deere, 2025

  • University of Arkansas, Agricultural Experiment Station

    United States, seven states

    Independent field research measuring targeted spraying against conventional broadcast application.

    An average yield increase of 2 bushels per acre, ranging up to 4.8 bushels per acre, against traditional broadcast spraying.

    University of Arkansas, 2025

  • Wageningen University and Research, with five international AI teams

    Netherlands

    A controlled comparison in which five teams set greenhouse climate, irrigation, lighting and CO2 remotely using their own algorithms, against a sixth compartment run by experienced Dutch growers, over a four month cucumber crop.

    Six identical 96 square metre compartments, August to December 2018. The winning team, Sonoma, out-produced the reference growers and reached a net profit of 24.78 euros per square metre against 21.18 euros for the growers, with the lowest CO2, heat and water use per kilogram of cucumber. The growers used the least electricity. The other four AI teams did not beat the growers on profit, and one used 13.61 kWh of heat per kilogram against the growers' 3.20 kWh.

    What it does not prove: One AI team out of five beat the human growers. For every team except one, crop handling decisions such as leaf pruning were set by expert policy rather than by the algorithms, and the authors note their training data did not cover pests and disease. A single four month crop in a research facility is not a commercial season.

    Sensors (MDPI), Hemming and colleagues, 2019

  • KU Leuven, the International Institute of Tropical Agriculture and CIMMYT

    Nigeria, northern maize belt

    A three year randomised controlled trial of Nutrient Expert, a tablet-based decision support tool that generates site-specific fertiliser recommendations, tested against the standard blanket recommendation of 120 kg nitrogen per hectare.

    792 smallholder maize farmers, 2016 to 2018. Site-specific advice on its own raised yield by about 8 per cent (roughly 173 kg per hectare) with no significant change in fertiliser use or revenue. The same advice paired with information on price variability and the distribution of returns raised yield by about 18 per cent (roughly 380 kg per hectare) and revenue by about 14 per cent, but also raised nitrogen use by about 18 per cent and greenhouse gas emissions per hectare by a similar margin. Nutrient use efficiency did not improve in either arm.

    What it does not prove: Nutrient Expert is a rule-based agronomic decision tool, not a machine learning system, so this measures the value of site-specific digital advice rather than of AI specifically. The headline gain only appeared when the tool was paired with risk and price information; the tool alone moved yield but not income, and the emissions result went the wrong way.

    Food Policy, via PubMed Central, 2023

  • Beach and colleagues, pooling 20 independent evaluations

    Sub-Saharan Africa, India and Cambodia

    A meta-analysis of 20 rigorous evaluations of digital information and advisory interventions delivered to small-scale farmers, covering interventions run between 2005 and 2019.

    Pooled effects were plus 23 per cent on fertiliser adoption (95 per cent confidence interval 6 to 40), plus 6 per cent on yield (CI 2 to 9) and plus 6 per cent on income (CI 2 to 9). Improved seed adoption at plus 11 per cent was not statistically significant (CI minus 6 to 28). Beneath the pooled figures, seven of the 13 studies measuring yield and five of the 10 measuring income found no statistically significant effect at all.

    What it does not prove: Covers digital advisory broadly rather than AI specifically, and the underlying interventions predate current models. The authors flag geographic concentration in eight countries, 15 of the 20 studies from Sub-Saharan Africa, thin cost-effectiveness evidence and possible publication bias towards positive findings.

    Global Food Security, via PubMed Central, 2025

  • University of Melbourne, CSIRO Data61 and the Singapore Food Agency

    Australia, Victoria, New South Wales and South Australia

    Random forest models predicting wheat yield from satellite vegetation index time series and climate records, evaluated both as a regional composite and separately on individual commercial paddocks.

    The three-paddock composite regional model reached R2 of 0.86 with RMSE of 0.18 tonnes per hectare. Paddock-level results split sharply: Victoria R2 0.89 (RMSE 0.15 t/ha) and New South Wales R2 0.87 (RMSE 0.07 t/ha), but South Australia only R2 0.45 (RMSE 0.25 t/ha). Across all models, high yields were systematically under-predicted and low yields over-predicted.

    What it does not prove: Retrospective modelling on a small number of paddocks, not an operational deployment. The authors attribute the weak South Australian result to soil variability within the paddock and to post-flowering drought and heat stress, which is precisely the season in which a grower would most want a reliable forecast.

    Sensors (MDPI), Pang, Chang and Chen, 2022

  • CSIRO

    Australia, South Australia

    Automated virtual fencing, in which GPS neckbands deliver an audio cue and then a mild electrical pulse based on the animal's position relative to a mapped boundary, used to keep cattle out of a regenerating river red gum area on a commercial farm.

    Twenty Santa Gertrudis heifers on a 14 hectare paddock over 44 days from May to July 2019, with a contoured (non-straight) virtual fence line. Cattle were excluded from the regenerating area for 99.8 per cent of the trial, spending one hour and 47 minutes inside it across 998 animal hours. Three of the 20 devices malfunctioned and were removed on day 39.

    What it does not prove: A small single-site trial of 20 animals over 44 days, run by the organisation that developed the underlying technology and licenced it commercially. The system is algorithmic geofencing combined with an animal-learning protocol rather than machine learning, so it belongs on an automation page more comfortably than on an AI one. Three of 20 collars failed during the trial, which is the number a farm operations manager will care about.

    Animals (MDPI), CSIRO authors, 2020

Back to all 13 sectors

Sector 11 of 13

Public sector and government

Not a sector we position in. Published here for the pattern.

Where the gains are real

The most transparent evidence available, because governments publish their costs. It is also the clearest working example of the pattern we design to: the machine does the volume, a named expert validates the output, and the validation time is counted in the total.

What the evidence does not support

The published figures compare against a manual process that was already slow. A saving against a bad baseline is real but says little about the ceiling.

  • UK Department for Science, Innovation and Technology

    United Kingdom

    Consult, part of the Humphrey toolset, grouping free-text public consultation responses into themes for expert review. Used on the Independent Water Commission's review of the water sector.

    More than 50,000 responses themed in around two hours at a compute cost of about 240 pounds, followed by 22 hours of expert human review and validation. The manual equivalent was weeks of sorting. Theme rankings were nearly identical to those produced by a team of human analysts.

    What it does not prove: The 22 hours of human validation is part of the result, not an overhead on it. Any comparison that omits it is not comparing the same thing.

    Think Digital Partners, 2025

  • UK Department for Science, Innovation and Technology

    United Kingdom

    Estimated the effect of applying the same approach across the government's annual consultation load.

    An estimated 75,000 person-days of analysis a year across more than 500 consultations, valued at roughly 20 million pounds in salary cost.

    What it does not prove: A departmental projection from one validated trial, not a realised saving.

    Global Government Forum, 2025

  • Australian Government Digital Transformation Agency

    Australia, Australian Public Service

    Whole of government trial of Microsoft 365 Copilot across the Australian Public Service, evaluated for the DTA by Nous Group using surveys, focus groups, interviews and agency internal evaluations.

    More than 2,000 trial participants from more than 50 agencies contributed to the evaluation of the 6 month trial announced in November 2023. 69 per cent agreed Copilot improved the speed at which they completed tasks and 61 per cent that it improved quality, with perceived savings of around an hour a day on summarising, first drafts and information searches. Only a third used Copilot daily, and up to 7 per cent reported that it added time to activities.

    What it does not prove: Every productivity figure is perceived and self-reported, not measured output. The evaluation itself states that Copilot's inaccuracy reduced the scale of the productivity benefits and that time spent verifying and editing outputs negated some of the efficiency savings.

    Digital Transformation Agency (digital.gov.au), 2024

  • Australian Department of the Treasury

    Australia, Commonwealth government

    Internal evaluation, conducted by the Australian Centre for Evaluation, of Treasury's Copilot trial, using pre-trial, pulse and post-trial surveys, a manager survey, focus groups, case studies and an issues log.

    218 staff took part over 14 weeks, from 20 May to 23 August 2024. Most participants reported using Copilot 2 to 3 times a week or less. Benefits were clearest for basic administrative tasks such as finding and summarising information and generating meeting minutes, but the evaluation found no clear evidence that Copilot improved work outcomes during the trial and did not explicitly measure time saved. It calculated that an APS6 officer would need to redirect roughly 13 minutes a week from low value to high value work to offset the licence cost.

    What it does not prove: A short trial in a restrictive IT security environment, which the evaluators say limited the product's performance relative to tools staff had used elsewhere. The 13 minute figure is a break-even calculation against licence cost, not a measured saving. Slight positive shifts in staff wellbeing could not be attributed to Copilot.

    Australian Centre for Evaluation, The Treasury, 2025

  • Government Technology Agency of Singapore (GovTech), Engineering Productivity Programme

    Singapore

    A 4 month evaluation of GitHub Copilot for Business across public sector software development teams, reported by GovTech's own engineering productivity team.

    70 developers signed up and 40 responded to surveys between October 2023 and January 2024. Respondents reported an average 22 per cent reduction in coding time, equivalent to a 28 per cent increase in coding speed, ranging from 33 per cent for junior developers to 15 per cent for seniors. Allowing for developers spending roughly half their time in the IDE, the authors put the overall productivity gain at about 12 per cent, or approximately 5 hours a week per developer.

    What it does not prove: Self-reported survey data from 40 respondents, published by the agency's own engineering team as a preprint without peer review. The authors themselves note the gains sit below the 30 to 100 per cent figures commonly quoted in industry, and attribute part of that to data security constraints on cloud-based tools.

    arXiv preprint by GovTech Singapore, 2024

  • Commonwealth of Pennsylvania, Office of Administration

    United States, Pennsylvania state government

    A year long pilot of ChatGPT Enterprise across state agencies, concluded in March 2025, with participants drawn from human resources, IT, policy and programme management roles.

    175 employees across 14 agencies were given licences and 136 provided direct feedback. 48 per cent had never used ChatGPT before, and 85 per cent reported a positive experience. Users estimated saving 95 minutes a day. The report's own finding list includes that ChatGPT is not a substitute for the nuance and experience employees bring to their work.

    What it does not prove: The 95 minutes is explicitly an estimate by volunteer participants, not a measured saving. The report is co-branded with the vendor, OpenAI, and states that the pilot population is not representative of the Commonwealth's workforce at large.

    Commonwealth of Pennsylvania, Office of Administration, 2025

Back to all 13 sectors

Sector 12 of 13

Telecommunications and media

Not a sector we position in. Published here for the pattern.

Where the gains are real

Australia has one of the better-documented deployments in the world here, and the interesting number is not handle time. It is the reduction in customers having to make contact a second time, which is the metric a service business actually loses money on.

What the evidence does not support

The strongest figures are agent-reported perception rather than measured handle time, and they are published by the technology vendor.

  • Telstra

    Australia

    Two tools for the service desk: One Sentence Summary, consolidating a customer's interaction history into a short brief, and Ask Telstra, natural-language retrieval over internal knowledge.

    20 per cent less follow-up contact. 90 per cent of employees using One Sentence Summary reported time savings and greater effectiveness. 84 per cent of Ask Telstra users agreed it improved customer interactions.

    What it does not prove: Vendor-published with the customer named and consenting. The 90 and 84 per cent figures are agent survey responses. The 20 per cent follow-up reduction is the operational one.

    Microsoft customer story, 2024

  • Australian Bureau of Statistics

    Australia

    Measured AI adoption by industry across Australian businesses.

    Information, media and telecommunications led every Australian industry at 38 per cent adoption, well ahead of professional services and financial services on 24 per cent.

    Australian Bureau of Statistics, 2026

  • Vodafone Group

    Europe, initially Italy and Portugal

    SuperTOBi, a generative AI virtual assistant replacing the earlier rule-based TOBi chatbot in customer service.

    Vodafone reported that in its first markets SuperTOBi "increased the first-time resolution rate from 15% to 60%" and that the online net promoter score "improved by 14 points to 64 points". The rollout was backed by EUR 140 million of investment in that financial year, with Germany and Turkey following Italy and Portugal.

    What it does not prove: Company-published and not audited. The 15 per cent baseline is Vodafone's own measure of its previous chatbot, so the improvement partly reflects how weak the starting point was. The figures cover the first markets only, not the group.

    Vodafone Group newsroom, 2024

  • Verizon and Google Cloud

    United States

    A "Personal Research Assistant" built on Google Cloud's Gemini models, giving frontline customer care staff answers drawn from Verizon's knowledge bases instead of manual searching.

    The joint announcement states the assistant reached "28,000 of Verizon's customer care reps and retail stores" and delivers "95% comprehensive answerability for customer inquiries".

    What it does not prove: A joint vendor and customer announcement. "Answerability" measures whether the assistant produced an answer, not whether the customer's problem was resolved. The release claims improved resolution time for new representatives but publishes no figure for it, and no call handling time or repeat contact data is given.

    Google Cloud Press Corner, 2025

  • European Broadcasting Union and BBC (News Integrity in AI Assistants)

    18 countries, 14 languages

    A coordinated evaluation in which journalists at 22 public service media organisations assessed AI assistant answers to news questions for accuracy, sourcing and the separation of fact from opinion.

    Across more than 3,000 assessed responses, 45 per cent had at least one significant issue, 31 per cent had serious sourcing problems including missing or misattributed citations, and 20 per cent contained significant accuracy errors such as fabricated or outdated detail.

    What it does not prove: The assessors are journalists at the broadcasters whose work the assistants were summarising, and those organisations have a commercial interest in the finding. The scoring rubric is published, but the judgements are human and not blinded. It remains the largest exercise of its kind and the per-assistant spread was wide.

    European Broadcasting Union, 2025

  • Deutsche Telekom

    Germany and group operations

    Group-wide internal AI tooling, including the askT employee assistant and the Frag Magenta customer assistant.

    In the annual report Deutsche Telekom states that in its November 2025 employee survey "53 % of our employees say they regularly use AI in their work, an increase of 9 percentage points against the survey from May 2025", with the proportion not using AI regularly down 8 percentage points.

    What it does not prove: Self-reported in an internal employee survey and published by the company. It measures how many people say they use AI, not what the use produced. Deutsche Telekom publishes no time saved, cost or resolution figures in this section.

    Deutsche Telekom Annual Report 2025, 2025

  • Reuters Institute for the Study of Journalism, University of Oxford

    48 markets

    The Digital News Report's annual survey of how people find and trust news, including use of AI chatbots as a news source.

    The 2026 report finds 10 per cent of people use AI chatbots for news weekly, up from 7 per cent the previous year, rising to 16 per cent of under-35s. Trust in answers from AI chatbots sits at 20 per cent globally against 37 per cent for news overall. Growth was uneven: usage doubled year on year in South Korea, Greece and Spain, while the United States and United Kingdom reported no increase.

    What it does not prove: An online survey, so it under-represents people with limited internet access and older age groups in some markets. Behaviour is self-reported rather than observed.

    Reuters Institute Digital News Report 2026, 2026

Back to all 13 sectors

Sector 13 of 13

Technology and software engineering

Not a sector we position in. Published here for the pattern.

Where the gains are real

The most-studied sector, and the one where the evidence most obviously contradicts itself. Controlled task studies show large speed-ups. Field studies on real repositories with experienced maintainers show the opposite. Both are properly conducted.

What the evidence does not support

The contradiction is the finding. Gains appear on unfamiliar, well-specified, greenfield work and disappear on complex code the developer already knows. Any vendor quoting only the first number is quoting half the literature.

  • GitHub and Microsoft Research

    Controlled experiment

    A randomised controlled trial asking developers to implement an HTTP server in JavaScript, with and without an AI coding assistant.

    The assisted group completed the task 55.8 per cent faster.

    What it does not prove: One well-specified greenfield task with a clear finish line. This is the most favourable possible test for AI assistance.

    arXiv, 2023

  • METR

    Field experiment on real repositories

    A randomised controlled trial with 16 experienced open-source maintainers across 246 real tasks in repositories they already knew well.

    Work took 19 per cent longer with AI tools available. The same developers had forecast a 24 per cent speed-up and still believed afterwards they had been 20 per cent faster.

    What it does not prove: 16 developers is a small sample, and the tooling generation moves quickly. The perception gap is the durable finding.

    METR, 2025

  • Stanford HAI, AI Index

    Global

    Aggregated measured productivity effects across business functions.

    About 26 per cent in software development, sitting between customer support at 14 to 15 per cent and marketing output at the top of the range.

    Stanford HAI, AI Index Report, 2026

  • Google (Paradis, Grey, Madison, Nam, Macvean, Meimand, Zhang, Ferrari-Church and Chandra)

    Peer-reviewed research, United States

    A randomised controlled trial inside Google in which engineers completed a realistic enterprise coding task, integrated with Google's own build and test systems, with and without internal AI tooling.

    Across 96 full-time Google software engineers, the AI-assisted group completed the task about 21 per cent faster. The authors state plainly that the confidence interval is large and that the result may not generalise beyond the tools and period studied (summer 2024).

    What it does not prove: A single task at a single company with 96 participants, and the authors themselves flag the wide confidence interval. It measures time on one well-scoped task, not throughput over a quarter, and says nothing about defects or review load downstream.

    arXiv (paper 2410.12944), 2024

  • Cui, Demirer, Jaffe, Musolff, Peng and Salz

    Peer-reviewed research, United States

    Three randomised field experiments giving GitHub Copilot to developers at Microsoft, Accenture and an anonymous Fortune 100 company, measured over months of normal work rather than a set task.

    Pooled across 4,867 developers, access to the assistant produced a 26.08 per cent increase in completed tasks, with a standard error of 10.3 per cent. Less experienced developers adopted the tool more and gained more from it.

    What it does not prove: The standard error of 10.3 per cent is large relative to the 26 per cent point estimate, so the true effect could be much smaller. Output is counted as completed pull requests, which measures work finished rather than value delivered or defects avoided, and two of the three sites are the tool vendor and a major implementation partner.

    Management Science, 2026

  • DORA (Google Cloud)

    Global survey

    The 2025 State of AI-assisted Software Development report, surveying software delivery practice and correlating AI adoption with delivery performance.

    Across nearly 5,000 technology professionals, 90 per cent reported using AI at work and more than 80 per cent believed it had increased their productivity, while 30 per cent reported little or no trust in the code AI generates. DORA reports that AI adoption continues to have a negative relationship with software delivery stability, even though its relationship with throughput turned positive this year.

    What it does not prove: A cross-sectional survey, so the relationships are correlations and cannot establish that AI causes instability. Productivity is self-reported by the same respondents who chose to adopt the tools, and the research is funded and published by a vendor with a commercial interest, though the methodology is public.

    Google Cloud blog (2025 DORA report announcement), 2025

  • Perry, Srivastava, Kumar and Boneh (Stanford University and UC San Diego)

    Peer-reviewed research, United States

    A controlled user study in which participants completed five security-related programming tasks across Python, JavaScript and C, with or without access to an AI code assistant based on OpenAI's codex-davinci-002.

    Among 47 participants (33 with the assistant, 14 without), those with access wrote insecure solutions more often on four of the five tasks after controlling for prior security exposure and programming experience. They were also more likely to believe their code was secure, so the quality gap came with more confidence rather than less.

    What it does not prove: A small sample of 47, largely students recruited at two universities, using a late 2022 model that has since been superseded. The finding that matters for a 2026 audience is the confidence gap rather than the specific vulnerability rates.

    ACM CCS 2023 (arXiv 2211.03622), 2023

  • Meta (Murali, Maddila, Ahmad, Bolin, Cheng, Ghorbani, Fernandez, Nagappan and Rigby)

    Peer-reviewed research, United States

    CodeCompose, Meta's internal AI code authoring assistant, fine-tuned on internal code and deployed across the engineering organisation, evaluated with both telemetry and developer surveys.

    At the time of writing, 16,000 developers had used CodeCompose with 8 per cent of their code coming directly from the tool. Fine-tuning on internal code let the model reproduce hidden lines between 40 and 58 per cent of the time, an improvement of 1.4 to 4.1 times over a model trained only on public data. Of the qualitative feedback collected, 91.5 per cent was positive.

    What it does not prove: Written by the team that built and owns the tool. The 8 per cent of code accepted is a usage measure, not a productivity or quality measure, and the 91.5 per cent positive feedback is voluntary self-reported sentiment from adopters rather than a controlled comparison.

    arXiv (paper 2305.12050), 2023

Back to all 13 sectors

What the evidence has in common

Six things every result on this page shares.

Read across all 13 sectors and the successful deployments look more like each other than they look like their own industries.

  • The task had a checkable right answer. Clinical notes, weeds, defects, consignment scans and consultation themes can all be verified by someone. Where correctness was a matter of opinion, the measured gains collapsed.
  • A named person stayed in the loop, and their time was counted in the total. The UK consultation result includes 22 hours of expert validation. Any business case that leaves review time out is not describing the same process.
  • The data existed before the project started. Rio Tinto had telemetry, Kaiser Permanente had structured encounters, John Deere had imagery. Nobody in this evidence base built a model first and found the data later.
  • The work was high-volume and repetitive. Every credible result comes from something done thousands of times, where a small per-instance saving compounds into a number worth reporting.
  • Somebody kept operating it. The published healthcare results describe continuous quality assurance over generated notes. The extraction results depend on new document formats being onboarded as they appear.
  • The escalation path was built before the automation was scaled. The one well-documented reversal in this evidence base happened where it was not.

Using this honestly

What a sector benchmark can and cannot tell you.

It can tell you what to try first

If document-heavy review work is where your sector’s evidence sits, that is a reasonable place to start looking. Evidence is a good guide to sequencing.

It cannot tell you what you will get

A 15,791-hour saving across 7,260 physicians says nothing about a 40-person business. Sample, baseline, data quality and starting maturity drive these numbers more than the technology does.

It cannot substitute for your own baseline

The only figure that will survive scrutiny inside your business is the one measured against your own process before anything changed. That measurement is the first thing worth funding.

Every figure on this page belongs to the organisation or study cited beside it. Links are to the primary source wherever one is published. Inclusion is not endorsement of any vendor, product or method, and Celestique Cloud publishes no outcome figures of its own. Sources last verified 30 July 2026.

Free discovery workshop

Work out which of these patterns applies to you.

Bring one process you think is a candidate. We will tell you whether the evidence supports it, what data it would need, and where it would have to stop for a person to approve something.