Skip to main content

Envisioning is an emerging technology research institute and advisory.

LinkedInInstagramGitHub

2011 — 2026

research
  • Observatory
  • Newsletter
  • Methodology
  • Origins
  • Vocab
services
  • Signals Session
  • Bespoke Projects
  • Build Sessions
  • Use Cases
  • Readinessfree
  • Signals
  • Free scan↗free
impact
  • ANBIMAFuture of Brazilian Capital Markets
  • IEEECharting the Energy Transition
  • Horizon 2045Future of Human and Planetary Security
  • WKOTechnology Scanning for Austria
solutions
  • Innovation
  • Strategy
  • Consultants
  • Foresight
  • Associations
  • Governments
resources
  • Partners
  • Coding for Non-Coders
  • How We Work
  • Data Visualization
  • Multi-Model Method
  • FAQ
  • Security & Privacy
about
  • Manifesto
  • Community
  • Events
  • Support
  • Contact
ResearchServicesSignalsAbout
ResearchServicesSignalsAbout
  1. Home
  2. Vocab
  3. DeepSeek R1

DeepSeek R1

A reasoning model trained with RL that showed emergent chain-of-thought without SFT.

Year: 2025Generality: 780Added: May 15, 2026
Back to Vocab

DeepSeek R1 is a large language model designed for extended reasoning tasks, developed by the Chinese AI lab DeepSeek. Unlike prior reasoning models that relied heavily on supervised fine-tuning, R1 was trained primarily through reinforcement learning, allowing it to develop chain-of-thought reasoning organically.

The model gained attention for matching or exceeding the performance of OpenAI's o1 on benchmarks including mathematics, coding, and scientific reasoning—while being trained at a fraction of the cost. R1's success demonstrated that extended thinking behaviors could emerge without explicit human-labeled reasoning traces.

R1's training used a group relative policy optimization approach, rewarding correct reasoning steps rather than final answers. This produced unusual behaviors including self-correction, backtracking, and unusually long chains of thought. The model also exhibited a "aha moment" during training where it spontaneously developed problem-checking behaviors.

The release of R1 sparked debate about the true source of reasoning capability in LLMs and the necessity of synthetic reasoning data. It also accelerated competition in reasoning models globally. Limitations include occasional hallucinations in long reasoning chains and performance drops on tasks outside formal reasoning domains.

Research this in Signals

Scan DeepSeek R1 for yourself.

Signals turns a topic into a sourced research record you can inspect and rerun. Your first scan is free, and this one starts with DeepSeek R1 already loaded, so edit it or scan as is.