The evidence note
Every external claim in the workshop, sorted by how much weight it can carry
Hey there.
Most AI statistics you meet in a slide deck have had their caveats surgically removed. The number survives. The sample size, the setting and the limitation do not. Six months later somebody repeats it in a board meeting as though it were a law of physics, and nobody in the room can say where it came from.
So here is the ledger behind the workshop. Every external claim I use, sorted into three classes by how much weight it can actually bear, with the limitation sitting next to the finding instead of in a footnote nobody reads.
Two things before the list.
The class matters more than the number. An independent randomised trial and a vendor-sponsored survey can report the same percentage and mean entirely different things. Most arguments about AI productivity are really arguments between evidence classes, and almost nobody says so out loud.
One entry changed while I was checking it. The most quoted study in this pack has been superseded by the people who ran it. I have left it in, with what they say now, because how that happened teaches more than the original finding ever did.
CLASS 1 | Independent controlled studies
Randomised or quasi-experimental, run by researchers without a product to sell. This is the only class that can support a claim about cause.
Cui, Demirer, Jaffe, Musolff, Peng and Salz, on software developers
"The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers." Management Science, 2025, DOI 10.1287/mnsc.2025.00535. Preprint via SSRN. Three randomised trials at Microsoft, Accenture and an anonymous Fortune 100 company, run by those companies in the ordinary course of business. 4,867 developers combined.
The finding: a 26.08% increase in completed tasks (standard error 10.3%) among developers given an AI coding assistant. Less experienced developers adopted it more and gained more.
Limitations: the individual experiments were noisy and the results varied between them, so the pooled figure is doing a lot of work. Software development only. It measures completed tasks, which is not quality and not downstream rework. A task completed and later corrected by someone else still counts as completed.
METR, on experienced open-source developers
Becker, Rush, Barnes and Rein, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity," July 2025. metr.org, arXiv:2507.09089. A randomised controlled trial: 16 experienced developers, 246 real issues on their own large and mature repositories, February to June 2025, with the tooling of the day.
The finding: tasks took 19% longer when AI tools were allowed. The developers had forecast a 24% speed-up beforehand. Afterwards, having been measurably slower, they still believed AI had made them roughly 20% faster.
And then this happened. In February 2026 METR published a continuation and put a banner on the original paper saying the results are out of date and no longer reflect the current impact of AI models on open-source developer productivity. The second study ran from August 2025 with 57 developers (10 of them from the original), 143 repositories and over 800 tasks. In METR's sign convention a negative number is a speed-up:
| Study | Estimate | Confidence interval |
|---|---|---|
| Original, early 2025 | +19% (slower) | +2% to +39% |
| Repeat developers, late 2025 | −18% (faster) | −38% to +9% |
| Newly recruited developers, late 2025 | −4% (faster) | −15% to +9% |
What METR says about their own second result: it is only very weak evidence of a speed-up. Both of those intervals cross zero. They are redesigning the experiment because developers increasingly decline to take part rather than work without AI, and because between 30% and 50% of participants reported avoiding tasks they did not want to attempt unassisted. Both effects push the estimate down, so METR's own reading is that the true gains are likely higher than these figures show.
How to use it. The perception gap is the durable finding, and it is the only part the workshop leans on. People who had been measurably slower were confident they had been faster. That is a fact about self-report, and nothing in the 2026 data touches it.
How not to use it. "A study proved AI makes developers slower." It does not say that today, and the people who ran it say so in a banner at the top of their own page. If you have been using the 19% figure, this is the week to stop.
Brynjolfsson, Li and Raymond, on customer support
"Generative AI at Work." NBER working paper 31161 (2023), published in the Quarterly Journal of Economics, 2025. 5,172 customer support agents at a Fortune 500 firm, staggered rollout plus a pilot randomised trial.
The finding: productivity, measured as issues resolved per hour, up 15% on average, with wide variation between workers. The least experienced and lowest-skilled improved in both speed and quality. The most experienced and highest-skilled saw small gains in speed and small declines in quality.
Limitations: one firm, one function, and a staggered rollout is quasi-experimental rather than a clean trial.
Why the second half of that finding matters: it is the most commonly dropped sentence in the whole literature. The same intervention helped the junior half of the team and slightly degraded the senior half's output. An average of 15% hides both.
International Center for Law & Economics, review of the evidence
Fruits and Stout, "AI, Productivity, and Labor Markets: A Review of the Empirical Evidence," 5 February 2026. laweconcenter.org.
The findings: task-completion time reductions of roughly 15% to over 50% across writing, customer support, software development, accounting, law and translation, often with quality gains, and consistently larger benefits for less experienced workers, which produces skill compression. On labour markets, aggregate indicators through 2024 and 2025 show limited disruption despite rapid adoption. Where effects appear they concentrate in entry-level segments of highly exposed occupations, with senior employment largely stable and adjustment happening through task reallocation rather than mass displacement.
Limitations: a review, not a primary study. Task-level findings do not aggregate to firm-level outcomes, which is the single most common error in this whole area. ICLE has stated policy positions, so read it as a well-sourced argument rather than a neutral survey.
CLASS 2 | Practitioner surveys
Self-reported. These tell you what people believe and what they say they do. They cannot tell you what works, and adoption is not effectiveness.
SANS 2026 AI Cybersecurity Report
536 practitioners and 57 senior leaders, global.
- 78% report active AI use in cybersecurity, up from roughly half the field a year earlier.
- 27% describe mature production deployment.
- 63% report significant shortcomings in AI-powered threat detection and response, up from 45% a year earlier.
Limitation: self-report throughout. The useful part is not any single figure, it is that adoption nearly doubled while the reported shortcoming rate rose with it. Capability and confidence are moving together, which is the opposite of what a maturing technology usually looks like.
Gravitee, State of AI Agent Security, 2026
750 senior technology leaders, UK and USA.
- Mean agent monitoring coverage 52%, so roughly half of production agents run without security oversight or logging. That is barely moved from 46.96% in December 2025, while the agent estate itself doubled.
- 19.7% of organisations secured all agents before going live, up from 13.6% in December 2025.
- Only 9.5% secure more than 81% of their deployed agents.
- 85% report no formal accountability for agent behaviour.
Limitation: vendor-sponsored, with a direct commercial interest in the finding that governance is lagging. Use it to illustrate that governance lags deployment. Never quote it as a measured rate.
Writer, with Workplace Intelligence, enterprise AI adoption survey, 2026
2,400 respondents, split evenly between C-suite executives and employees, fielded 17 December 2025 to 25 January 2026.
- 67% of executives believe their company has already suffered a leak or breach because an employee used an unapproved AI tool.
Limitations: vendor-sponsored, and that 67% is a belief, not incident data. Executives who believe a leak happened and executives who can evidence one are different populations.
A number to be careful with, because the draft of this ledger had it wrong. An earlier version of this page reported "35% of employees enter proprietary information into public tools" and credited it to this survey. Checking it before publication showed that is not what the 35% measures. In this study, 35% is the share of executives who are not very confident they could stop a rogue AI agent that started causing damage. The employee-side figure is that 29% admit to undermining their company's AI strategy, and the survey defines that broadly enough to cover entering company information into public tools, using unapproved tools, or simply declining to use AI at all. Those are three quite different behaviours wearing one percentage.
This is exactly the failure mode the workshop is about. A real number, from a real survey, attached to the wrong claim, and it survives because it sounds right and nobody opens the source.
CLASS 3 | Vendor case studies
The weakest class. Present them as an example of the class, never as evidence.
Legal and procurement AI vendors publish customer case studies claiming 50% to over 90% reductions in first-pass contract review time, usually alongside a return-on-investment study the vendor commissioned. Named customers, unverifiable methods, no control group, and a commercial interest in the result.
These are not worthless. They tell you what a vendor is willing to put its name to, and they are a reasonable source of hypotheses about where to look. They are not a reason to do anything.
CLAIMS THIS PACK WILL NOT MAKE
- No maturity distribution. The four-stage maturity frame is a diagnostic. Any percentage split across those stages that you see in a deck, including in mine, has no identified source. The stages are useful. The distribution is invented.
- No client metric. Nothing from any organisation I have worked with appears here. Not a cost reduction, not a time saving, not anything. If it has not been independently verified, it does not travel, however good it sounds.
- No universal capability claim. "AI can now do X" without naming the conditions is the specific error this workshop teaches people to catch. The conditions are the claim.
- No return-on-investment promise. Any multiple you have heard belongs to someone's financial model, not to this programme.
- No legal specifics without specialist review. Where a regulatory question comes up I will note it and route it rather than answer it. And I will always say whether I am describing what the law requires or what a company's own policy requires, because those get conflated constantly and only one of them can be appealed.
HOW TO USE THIS
Three questions, and they work on any AI claim you meet, including mine.
- Which class is it? Independent trial, self-reported survey, or vendor case study. If you cannot tell, treat it as the weakest one.
- What is the setting? Who, doing what work, with which tools, when. A finding without a setting is an advertisement.
- Is it still current? The METR entry above is the whole argument for asking. A good finding went stale in seven months and the authors said so loudly, and it is still circulating in its original form today.
Every entry on this page was re-verified against the primary source on 6 October 2026. If you are reading this much later than that, assume at least one thing here has moved, and check the two METR links first, because that is where movement shows up fastest.
Found an error? I would genuinely rather know. contact@aleodor.com
Keep on learning, keep on building.