AI
Stanford's AI Index runs 396 pages. Six findings survive translation to a 30-person nonprofit

By
Damon Stewart
We read the full 2026 AI Index, including the appendix. Most of it describes a world nonprofits do not live in. Six findings transfer.
At some point this year a board member sends you the Stanford AI Index, or more likely a LinkedIn summary of it, with a note that says some version of: what are we doing about this? The report is the most authoritative general AI research artifact of the year, it is 396 pages plus an appendix, and it is almost entirely about a world your organization does not live in. We read all of it, including the appendix, on August 5, 2026. This piece reports what actually transfers to a nonprofit with 30 staff and a budget between 1 million and 50 million dollars, and what does not.
Credit where it belongs: Susan Mernit made the core argument in April 2026, that the Index's headline numbers describe organizations nothing like a small nonprofit. What we add is the methodology work underneath that argument, because the appendix is where this report gets interesting.
The 88 percent, and where it comes from
The number your board member will quote is on page 9: 88 percent of organizations now use AI in at least one business function, up from 78 percent last year. Before it enters your planning, walk it back to the appendix on page 396. The survey behind it drew 1,993 respondents across 105 nations, fielded June 25 to July 29, 2025, so the figure is roughly a year old already. Thirty-eight percent of respondents work at organizations with more than 1 billion dollars in annual revenue, and responses are weighted by each country's contribution to global GDP. And an organization counts as an adopter if AI is used in at least one business function, which means one person in one department using one tool makes the whole organization a yes.
None of that is hidden, and none of it is a scandal. The report says about its own survey data, in section 4.3: "As with other survey-based data in this chapter, the results are self-reported and should be viewed as directional rather than comprehensive." Stanford discloses what the number is. The people quoting it at you skip the disclosure. That is the same pattern we documented with the nonprofit sector's own 92 percent statistic, and the discipline is the same: a billion-dollar-weighted global enterprise survey cannot describe your organization, and the report never claims it can.
The number that comes closest to describing you is buried on page 197, in Figure 4.3.6, and it is the most useful exhibit in the report. Among the smallest organizations surveyed, still under 100 million dollars in revenue, so an order of magnitude above most nonprofits: 9 percent are not using AI at all, 39 percent are experimenting, 22 percent are piloting, 25 percent are scaling, and 5 percent have fully scaled. Read that again the next time someone implies you are behind. In the friendliest possible comparison group, experimenting is the normal state and full deployment is a twentieth of the population.
What does not transfer
Most of the report's most-quoted numbers carry no operational content for your decisions, and saying so specifically is the point of a translation.
The investment figures do not transfer. Global corporate AI investment of 581.69 billion dollars describes the supply side of an industry, not the decision facing an organization weighing 400 dollars a month in software licenses. Its only use to you is as context for why the tools keep changing underneath you.
The productivity studies transfer less than they appear to. Every measured gain in Figure 4.4.27 comes from structured, high-volume, easily monitored work: customer support agents resolving 14 to 15 percent more tickets, developers completing 26 percent more pull requests, accountants showing 55 percent throughput gains. The report says it plainly: gains are largest "in structured, measurable work where outputs are easy to monitor." A nonprofit's expensive labor is program delivery, funder relationships, and judgment under ambiguity, which is the category the report identifies as showing weaker or negative effects. The counter-evidence sits in the same section, and we carry it with its caveat: the METR study found experienced developers got 19 percent slower with AI assistance, and the METR team has not been able to replicate that result in a later study. Most people who quote the 19 percent drop the second half.
The workforce findings do not transfer. The documented employment decline is real and specific: software developers aged 22 to 25, down close to 20 percent from the 2022 peak. The report's own framing is that "the evidence does not point to broad, uniform displacement." A 30-person nonprofit has no junior developer bench to shrink, and the organizations expecting headcount cuts are concentrated in service operations and engineering at billion-dollar firms. Reading that section as a warning about your program staff is a category error.
The capability benchmarks do not transfer at all. Olympiad gold medals and PhD-level science scores change nothing about your Monday, because frontier capability has been sufficient for the tasks a small nonprofit would automate for at least two years. The binding constraint was never the model.
The macro evidence, for what it is worth, refuses to agree with itself, and that is the finding. One survey of 6,000 executives found widespread adoption and minimal realized productivity gains. The Penn Wharton Budget Model puts AI's current contribution to total factor productivity at 0.01 percentage points, which the report's own table calls negligible. A 12,000-firm European study found a 4 percent labor productivity gain. Anyone quoting a single one of these has chosen a conclusion and gone looking for its citation.
We will also report two things the report gets wrong or leaves unresolved, because a translation you can trust has to include them. The Economy chapter highlight says generative AI is used at 70 percent of organizations while the body text and the underlying chart both say 79, a nine-point internal contradiction the report does not notice. And it carries two hallucination benchmarks that measure different things: a new accuracy benchmark where rates across 26 top models range from 22 to 94 percent, and a document-summarization leaderboard where the top models sit between 1.8 and 5.4 percent. Quoting either alone misleads. The first is about models handling contested beliefs, and its sharpest finding deserves quoting: when a false statement is presented as something the user believes, model accuracy collapses. The machine is politest exactly when you need it to argue.
The six findings that survive
First, the deployment-stage distribution. Experimenting is the normal state at organizations far bigger than yours. Nobody at your scale has to feel behind, and any pitch that opens with the 88 percent is selling urgency the data does not contain.
Second, the report's own disclaimer. "Directional rather than comprehensive" is a permission slip: discount every adoption statistic anyone shows you, including this report's, including ours.
Third, the selection criterion. Gains cluster in structured, monitorable work. Your organization has some: inquiry response, appointment scheduling, enrollment follow-up, donation processing, report assembly. That list, not a strategy document, is where an AI project starts, because it is the only place the evidence says the returns are reliable.
Fourth, the staff-preference finding almost nobody covers, from Figure 4.4.35: 46.1 percent of workers actively want AI to take over specific tasks, and the tasks they most want automated account for 1.3 percent of observed AI usage. The staff resistance most executives fear is probably misdiagnosed. The real gap is that nobody has asked staff which tasks they would gladly hand over.
Fifth, the risk numbers that argue for review rather than avoidance. Documented AI incidents rose to 362 in 2025 from 233 the year before, and the belief-sensitivity result above says confident wrong answers are a design property, not a bug being fixed next quarter. The operational answer is output review: a human who checks the citation, the eligibility rule, the number, before it ships.
Sixth, the 50-point gap. Seventy-three percent of AI experts expect a positive impact on how people do their jobs, against 23 percent of the public. Your board lives on one side of that gap and your staff on the other, which is why the same agenda item produces two different conversations. Knowing the gap is measured, and enormous, is the beginning of running that meeting well.
Six usable findings from 396 pages is not a criticism of the report, which is careful, well documented, and simply not written for you. It is what translation actually looks like. When the next authoritative number arrives in your inbox, the questions are the ones that worked here: who was sampled, how is it weighted, when was it fielded, and does the report itself tell you how seriously to take it. The answers are usually in the appendix, and almost nobody reads the appendix.


