Skip to main content

Envisioning is an emerging technology research institute and advisory.

LinkedInInstagramGitHub

2011 — 2026

research
  • Observatory
  • Newsletter
  • Methodology
  • Origins
  • Vocab
  • RSS Feeds
services
  • Signals Session
  • Bespoke Projects
  • Build Sessions
  • Pricing
  • Use Cases
  • Signals
  • Free scan↗free
impact
  • ANBIMAFuture of Brazilian Capital Markets
  • IEEECharting the Energy Transition
  • Horizon 2045Future of Human and Planetary Security
  • WKOTechnology Scanning for Austria
solutions
  • Innovation
  • Strategy
  • Consultants
  • Foresight
  • Associations
  • Governments
  • L&D
resources
  • Partners
  • Coding for Non-Coders
  • How We Work
  • Data Visualization
  • Multi-Model Method
  • FAQ
  • Security & Privacy
  • Public Sector
about
  • Manifesto
  • Community
  • Events
  • Support
  • Contact
ResearchServicesSignalsAbout
ResearchServicesSignalsAbout
  1. Home
  2. Vocab
  3. Illusion of Thinking

Illusion of Thinking

Apple's 2025 critique — reasoning models' accuracy collapses at certain puzzle complexities.

Year: 2025Generality: 550Added: Aug 1, 2026
Back to Vocab

Opening

"Illusion of Thinking" is the title of a 2025 paper from Apple Machine Learning Research that became the central scientific critique of large reasoning models in 2025–2026. The paper showed that frontier LRMs — including OpenAI's o3 — exhibit "complete accuracy collapse" under surprisingly simple conditions: as puzzle complexity (measured by the number of moving pieces in classic logic puzzles like Tower of Hanoi, river crossing, and block stacking) increases past a model-specific threshold, accuracy drops from high to near-zero, then in some cases partially recovers at higher complexities in a way that suggests the models are recognizing puzzle patterns rather than reasoning about them.

Mechanism

The paper tested frontier LRMs on a family of classic puzzles with controllable complexity parameters. As complexity scaled, three regimes emerged for every model tested. In the low-complexity regime, models solved puzzles reliably. In a middle regime, accuracy degraded in ways that suggested the model was attempting to reason but hitting a step-following ceiling. In the high-complexity regime, accuracy collapsed to zero even though the puzzle remained formally solvable — sometimes even easy for a human with a notebook. The paper coined the operational term "inference horizon" for the complexity threshold past which the model can no longer follow a reasoning chain, and argued that this horizon is the central practical limit of LRM capability rather than the "emergent reasoning" narrative that LRM marketing tends to emphasize.

Tradeoffs

The paper's empirical findings are robust, but the interpretation is contested. Pro-LRM researchers (notably Sébastien Bubeck at OpenAI) argue that the negative results reflect a training-data quirk in models that are now obsolete — that GPT-5.5 and successors have pushed the inference horizon further and the 2025 results no longer apply. Critics of the critique argue that the paper's complexity parameters are artificial and the results conflate task difficulty with reasoning failure. The trade-off the paper surfaces is between taking negative empirical results seriously as a theory of model capability and dismissing them as a snapshot of model-version-specific failure modes that will be engineered away in the next release cycle.

Open Questions

Whether the "inference horizon" is a fundamental architectural limit of transformer-based LRMs or a training-time artifact that improves with scale and data. Whether the high-complexity-regime collapse pattern generalizes beyond the specific puzzles Apple tested (Tower of Hanoi and family) to other reasoning tasks. Whether newer LRM releases (GPT-5.5, Claude with extended thinking v2, Gemini Thinking 2) have demonstrably pushed the inference horizon or simply shifted which puzzles trigger the collapse. Whether the paper's framing has been a net positive for the field (raising the bar for "reasoning" claims) or a net negative (distracting from the verifiable-domain successes that LRMs do deliver).

Research this in Signals

Scan Illusion of Thinking for yourself.

Signals turns a topic into a sourced research record you can inspect and rerun. Your first scan is free, and this one starts with Illusion of Thinking already loaded, so edit it or scan as is.