---
title: Illusion of Thinking
type: vocabulary
url: "https://www.envisioning.com/vocab/illusion-of-thinking"
summary: "Apple's 2025 critique — reasoning models' accuracy collapses at certain puzzle complexities."
year: 2025
generality: 0.55
---

# Illusion of Thinking

Apple's 2025 critique — reasoning models' accuracy collapses at certain puzzle complexities.
## Opening
"Illusion of Thinking" is the title of a 2025 paper from Apple Machine Learning Research that became the central scientific critique of large reasoning models in 2025–2026. The paper showed that frontier LRMs — including OpenAI's o3 — exhibit "complete accuracy collapse" under surprisingly simple conditions: as puzzle complexity (measured by the number of moving pieces in classic logic puzzles like Tower of Hanoi, river crossing, and block stacking) increases past a model-specific threshold, accuracy drops from high to near-zero, then in some cases partially recovers at higher complexities in a way that suggests the models are recognizing puzzle patterns rather than reasoning about them.

## Mechanism
The paper tested frontier LRMs on a family of classic puzzles with controllable complexity parameters. As complexity scaled, three regimes emerged for every model tested. In the low-complexity regime, models solved puzzles reliably. In a middle regime, accuracy degraded in ways that suggested the model was attempting to reason but hitting a step-following ceiling. In the high-complexity regime, accuracy collapsed to zero even though the puzzle remained formally solvable — sometimes even easy for a human with a notebook. The paper coined the operational term "inference horizon" for the complexity threshold past which the model can no longer follow a reasoning chain, and argued that this horizon is the central practical limit of LRM capability rather than the "emergent reasoning" narrative that LRM marketing tends to emphasize.

## Tradeoffs
The paper's empirical findings are robust, but the interpretation is contested. Pro-LRM researchers (notably Sébastien Bubeck at OpenAI) argue that the negative results reflect a training-data quirk in models that are now obsolete — that GPT-5.5 and successors have pushed the inference horizon further and the 2025 results no longer apply. Critics of the critique argue that the paper's complexity parameters are artificial and the results conflate task difficulty with reasoning failure. The trade-off the paper surfaces is between taking negative empirical results seriously as a theory of model capability and dismissing them as a snapshot of model-version-specific failure modes that will be engineered away in the next release cycle.

## Open Questions
Whether the "inference horizon" is a fundamental architectural limit of transformer-based LRMs or a training-time artifact that improves with scale and data. Whether the high-complexity-regime collapse pattern generalizes beyond the specific puzzles Apple tested (Tower of Hanoi and family) to other reasoning tasks. Whether newer LRM releases (GPT-5.5, Claude with extended thinking v2, Gemini Thinking 2) have demonstrably pushed the inference horizon or simply shifted which puzzles trigger the collapse. Whether the paper's framing has been a net positive for the field (raising the bar for "reasoning" claims) or a net negative (distracting from the verifiable-domain successes that LRMs do deliver).

---
Source: Envisioning — Technology Research Institute (https://www.envisioning.com/vocab/illusion-of-thinking)
