cursor
S U R E S H   M A N I C K A M
AI & Tools

The 2026 AI Index in Plain English: 10 Numbers Worth Your Time

· 6 min read

The 2026 AI Index in Plain English: 10 Numbers Worth Your Time

Every year Stanford's Institute for Human-Centered AI publishes the AI Index, and every year it does the same quietly valuable thing: it replaces vibes with numbers. The 2026 edition is out, it runs across nine chapters, and it is free. If you only read one AI document this year, make it this one.

I went through it looking for the things that change how I plan, budget and argue at work, not the things that make a good headline. Here are ten of them, and then the parts that deserve a bit of caution.

Read it yourself: The 2026 AI Index Report · Full report (PDF)

1. The money more than doubled

Global corporate AI investment more than doubled in 2025. Private investment alone grew 127.5 percent and now makes up 60 percent of the total. Generative AI grew more than 200 percent and took nearly half of all private AI funding. Newly funded AI companies rose 71 percent, and billion-dollar funding rounds nearly doubled. Whatever you think of the valuations, the capital is real and it is still arriving.

2. Adoption is near-universal, agents are not

88 percent of surveyed organisations now report using AI, and 70 percent use generative AI in at least one business function. But AI agent deployment sits in the single digits across nearly every business function. That gap is the whole story of 2026. Everyone has a chatbot. Almost nobody has agents doing real work in production yet.

3. Generative AI spread faster than the internet did

53 percent adoption in three years, quicker than the personal computer or the internet at the same stage. The surprise is where. Adoption tracks GDP per capita, but Singapore at 61 percent and the UAE at 64 percent both beat what their income would predict, while the United States, the country that builds most of the models, ranks 24th at 28.3 percent. Building the technology and using it are apparently two different national skills.

4. The performance gap between the top labs has almost vanished

As of March 2026, Anthropic, xAI, Google and OpenAI sit within 25 Elo points of each other on the Arena leaderboard, with Alibaba and DeepSeek just behind. The U.S. and China model gap has effectively closed too, with the top U.S. model leading by 2.7 percent after the two traded places repeatedly through the year. When raw capability converges, the competition moves to cost, reliability and how well a model does your specific job. That is a much better world for buyers.

5. Benchmarks are being eaten faster than they can be built

Frontier models gained 30 percentage points in a single year on Humanity's Last Exam, a test deliberately designed to be brutal for AI. Evaluations meant to last years are being saturated in months. Worse, the benchmarks themselves are shaky: a review found invalid question rates ranging from 2 percent on MMLU Math to 42 percent on GSM8K. Be careful how much weight you put on a leaderboard.

6. Intelligence is jagged, and it is funny about it

Google's Gemini Deep Think scored a gold medal at the 2025 International Mathematical Olympiad, working end to end in natural language inside the 4.5-hour limit. The same class of model reads an analog clock correctly 50.6 percent of the time, against 90.1 percent for humans. Olympiad maths, yes. Telling the time, not so much. Never assume that because a model can do the hard thing, it can do the easy one.

7. Agents got real, and still fail one attempt in three

On OSWorld, which tests agents on genuine computer tasks across operating systems, accuracy jumped from roughly 12 percent to 66.3 percent, within six points of human performance. That is a remarkable year. It is also a failure rate you cannot put in front of a customer without a human in the loop. Design for the third of the time it gets it wrong.

8. The productivity gains are real but narrow

Studies report 14 to 15 percent gains in customer support, 26 percent in software development, and 50 percent in marketing output. All of those are structured, measurable work where you can see the output. Gains shrink on tasks needing deeper reasoning, and the Index flags early evidence that heavy reliance on AI may carry a long-term learning penalty that slows skill development. Speed now, skill debt later, is a trade worth watching in your own team.

9. The labour effect is showing up at the entry level first

Employment for software developers aged 22 to 25 has fallen nearly 20 percent from 2024. One third of organisations expect AI to reduce their workforce in the coming year, with the largest expected cuts in service operations, supply chain and software engineering. Overall employment data has not moved much yet. The pipeline has. If you care about building a bench of talent, this is your problem to solve, not HR's.

10. Transparency went backwards

This is the one that bothers me most. The average score on the Foundation Model Transparency Index rose from 37 to 58 between 2023 and 2024, then dropped to 40 in 2025. Disclosure about training data, compute and post-deployment impact all got worse. Meanwhile documented AI incidents rose to 362 in 2025, up from 233. More incidents, less information. That is the wrong direction for anyone doing vendor due diligence.

The parts that deserve caution

  • Models confuse knowledge with belief. Hallucination rates across 26 top models range from 22 to 94 percent on a new accuracy benchmark. Tell a model that someone else believes a false thing and it handles it fine. Tell it that you believe it, and performance collapses. GPT-4o dropped from 98.2 to 64.4 percent, DeepSeek R1 from over 90 to 14.4 percent. Sycophancy is a measurable failure mode now.
  • Safety holds up until someone pushes. Frontier models score well on the AILuminate safety benchmark under normal use, then degrade across the board under adversarial prompts. Your red team matters more than the vendor's safety card.
  • English still gets the best model. On a Slovenian commonsense reasoning test, leading models lost close to half their accuracy when the question was posed in a regional dialect. If you serve a multilingual market, test in the language your customers actually speak.
  • Responsible AI goals fight each other. Empirical studies found that training to improve one dimension, say fairness, consistently degraded another, say privacy or safety. There is no free lunch and no single dial to turn.

What people actually think

The public opinion chapter is the one I would put in front of a leadership team. Globally, 59 percent now say AI offers more benefits than drawbacks, up from 55 percent, while 52 percent say it makes them nervous. Both numbers went up at once, which is a very human place to be.

The expert gap is stark. 73 percent of AI experts expect a positive impact on how people do their jobs. Only 23 percent of the U.S. public agrees. That is a fifty point gap on the single question your staff care about most. If you are rolling out AI internally and communicating from the expert side of that gap, you are talking past most of the room.

One more that stood out from my part of the world: more than 80 percent of respondents in Malaysia, Thailand, Indonesia and Singapore expect AI to profoundly change their lives in the next three to five years, and Singapore records the highest trust in its own government to regulate AI at 81 percent. The United States reports the lowest, at 31 percent.

What I am taking away

The Index does not tell you AI is over-hyped or under-hyped. It tells you something more useful: capability is arriving faster than measurement, adoption is arriving faster than governance, and the benefits are concentrating in work that is easy to observe. All three of those are things a technology leader can act on.

My short list. Stop buying on benchmarks and start running your own evaluations on your own tasks. Assume agents fail a third of the time and design the human checkpoint before you design the workflow. Ask vendors the transparency questions the Index says they have stopped answering voluntarily. And put real effort into your junior pipeline, because the data says that is where the squeeze is landing first.

Then go and read the thing. It is 300-odd pages of evidence, it is free, and it will make you better at every AI conversation you have this year.

The 2026 AI Index Report, Stanford HAI · Download the full PDF