Overview
Report
Coverage
Explorer
Sector Intelligence Report
State of AI: Reasoning Models
The global landscape, the deployment evidence, and the decision framework for 2026.
Swipe to begin ›Report 2026-026
01
Finding 01 · a new class
A new class
Reasoning models check their own logic before answering, which makes them measurably more accurate on high-value work. That also means their business case must be judged use case by use case, not folded into an existing AI programme.
evaluate them separately
Next finding ›Skip to Coverage
02
Finding 02 · the landscape
6 tiers
The market restructured into six vendor groups. The largest — thirteen China-headquartered vendors — matches or leads Western firms on four of five standard benchmarks at five-to-forty times lower cost.
the cheapest now match the best
Next finding ›Skip to Coverage
03
Finding 03 · the sequence
Admissibility first
Which vendors you are permitted to use — data residency, compliance, geopolitics — is a legal question, settled before any benchmark. Opening a shortlist first answers that legal question in the wrong room.
a legal decision, not a benchmark one
Next finding ›Skip to Coverage
04
Finding 04 · the gap
95%
of enterprise AI pilots produce no measurable P&L impact (EY, MIT Technology Review). Analyst forecasts run far ahead of production reality, so in the eight evidence-thin use cases the pilot is the wrong instrument without a hard evidence gate.
forecasts run ahead of evidence
The plan ›Skip to Coverage
The debrief · 3 minutes

The decision framework for reasoning models in 2026

This is AG Insights' sector read on reasoning models: forty-plus vendors across six tiers, the deployment evidence in twelve use cases, and the two-question framework that turns "should we adopt AI?" into a grounded, use-case-by-use-case decision. Start with the short audio debrief; the full analysis, the exhibits, and every underlying source sit in the tabs alongside.

4 use cases mature · 6 vendor tiers · 12,800+ sources screened
Listen to the debrief · the partner's readout, ~3 min
A spoken "this is what we found and how to decide" — the fast way in before the report.
What we did

Profiled 43 reasoning-model vendors across the evidence dimensions that matter — flagship model, weights, pricing, benchmarks and deployment — reading each firm's own material, then triangulated the deployment record in twelve use cases and the projections of eighteen analyst institutions against independent sources.

What we found
  • Reasoning models are mature in four use cases and near-null in the other eight; the split is the decision.
  • The landscape is six tiers; the largest, thirteen Chinese vendors, matches the West on four of five benchmarks at 5–40× lower cost.
  • Admissibility precedes capability — which tiers you may use is a legal question, settled before benchmarks.
  • 95% of enterprise AI pilots produce no measurable P&L; forecasts run ahead of production evidence.
What we recommend
  1. Write a tier policy first — before any shortlist opens.
  2. Gate on named-deployment evidence in the four rich use cases.
  3. Redesign the pilot in the eight thin ones — P&L gate and hard exit.
  4. Make the gate a board commitment — pre-commit to exit.
Settle admissibility, then evaluate capability.
Read the full report →
Coverage · every entity × dimension, coloured by confidence
Verified
Reliable
Indicative
Thin / gap
Three datasets — vendors, analyst institutions, and enterprise use cases. Colour is confidence, the number is sources on file. Click any cell to open it in the Explorer.