Skip to main content

Envisioning is an emerging technology research institute and advisory.

LinkedInInstagramGitHub

Since 2010

research
  • Observatory
  • Adaptive capacity
  • Newsletter
  • Methodology
  • Origins
  • Vocab
  • RSS feeds
services
  • Signals Session
  • Bespoke Projects
  • Build Sessions
  • Pricing
  • Use cases
  • Signals
  • Signal Scan↗free
impact
  • ANBIMAFuture of Brazilian Capital Markets
  • IEEECharting the Energy Transition
  • Horizon 2045Future of Human and Planetary Security
  • WKOTechnology Scanning for Austria
solutions
  • Innovation
  • Strategy
  • Consultants
  • Foresight
  • Associations
  • Governments
  • L&D
resources
  • Partners
  • Coding for Non-Coders
  • How we work
  • Data visualization
  • Multi-Model Convergence
  • FAQ
  • Security and privacy
  • Public sector
about
  • Manifesto
  • Community
  • Events
  • Support
  • Contact
ResearchCapabilityServicesSignalsAbout
ResearchCapabilityServicesSignalsAbout
  1. Home
  2. Vocab
  3. Scalable Oversight

Scalable Oversight

Methods for supervising AI systems whose outputs exceed human ability to evaluate directly, including debate, recursive reward modeling, and weak-to-strong generalization.

Year: 2018Generality: 500Added: Sep 4, 2026
Back to Vocab

Scalable oversight is the problem of supervising AI systems that produce outputs humans cannot reliably evaluate. The setting arises when models become more capable than the humans rating their outputs: a human annotator cannot judge whether a long, expert-level answer is correct, but the answer still needs to be scored for training.

Proposed approaches include AI debate (two models argue opposing positions, a human judge picks the winner), recursive reward modeling (a smaller model rates a larger model's output, the smaller model is itself rated by an even smaller model), weak-to-strong generalization (using a weaker model's labels to train a stronger one), and constitution-based methods (rule-following rather than preference scoring).

A 2026 paper by Bergemann, Koh, and Morris treats scalable oversight as a mechanism design problem: the question is what contracts can elicit honest work from an agent whose capability exceeds the principal's ability to verify the work directly.

Sources

  1. Mechanism Design for Alignment and Control

    arXiv · Sep 1, 2026

  2. AI safety via debate (Irving, Christiano, Amodei)

    arXiv · May 2, 2018

  3. Weak-to-strong generalization (OpenAI)

    OpenAI · Dec 6, 2023

Research this in Signals

Scan Scalable Oversight for yourself.

Signals turns a topic into a sourced research record you can inspect and rerun. Your first scan is free, and this one starts with Scalable Oversight already loaded, so edit it or scan as is.