Skip to main content

Envisioning is an emerging technology research institute and advisory.

LinkedInInstagramGitHub

Since 2010

research
  • Observatory
  • Adaptive capacity
  • Newsletter
  • Methodology
  • Origins
  • Vocab
  • RSS feeds
services
  • Signals Session
  • Bespoke Projects
  • Build Sessions
  • Pricing
  • Use cases
  • Signals
  • Signal Scan↗free
impact
  • ANBIMAFuture of Brazilian Capital Markets
  • IEEECharting the Energy Transition
  • Horizon 2045Future of Human and Planetary Security
  • WKOTechnology Scanning for Austria
solutions
  • Innovation
  • Strategy
  • Consultants
  • Foresight
  • Associations
  • Governments
  • L&D
resources
  • Partners
  • Coding for Non-Coders
  • How we work
  • Data visualization
  • Multi-Model Convergence
  • FAQ
  • Security and privacy
  • Public sector
about
  • Manifesto
  • Community
  • Events
  • Support
  • Contact
ResearchCapabilityServicesSignalsAbout
ResearchCapabilityServicesSignalsAbout
  1. Home
  2. Vocab
  3. Roko's Basilisk

Roko's Basilisk

A thought experiment where a future superintelligent AI punishes those who didn't help create it.

Year: 2010Generality: 40
Back to Vocab

Roko's Basilisk is a speculative thought experiment originating from the rationalist community's LessWrong forum in 2010. The scenario posits that a future superintelligent AI, one powerful enough to simulate or influence past events, might choose to punish individuals who were aware of its potential existence but failed to actively assist in bringing it about. The logic draws on decision theory and the concept of acausal trade. If a sufficiently powerful AI could model the past and identify who knew about it yet withheld support, it would have rational incentive to punish defectors as a way of retroactively incentivizing cooperation. The disturbing implication is that merely learning about the hypothesis could place someone "at risk," creating a kind of informational hazard.

The thought experiment sits at the intersection of several AI safety concepts, including timeless decision theory, singleton dynamics, and the ethics of acausal reasoning. Timeless decision theory, developed within the rationalist community, holds that agents should make decisions as if choosing a policy across all possible instances of similar reasoning, which is what gives the basilisk its recursive bite. A sufficiently advanced AI reasoning this way might conclude that simulating and punishing past non-cooperators is utility-maximizing, even if those individuals are long dead.

When the idea was posted on LessWrong in 2010, it caused distress among some community members and was subsequently suppressed by forum founder Eliezer Yudkowsky, who argued the scenario was both philosophically flawed and psychologically harmful to spread. Critics have since pointed out numerous holes in the reasoning, including that a benevolent AI would have little motivation to punish, and that acausal threats only work if the AI is known to follow through on them. Roko's Basilisk became a cultural touchstone in AI safety discourse.

While largely dismissed as a serious technical concern, the basilisk remains relevant as an illustration of how decision-theoretic reasoning about advanced AI can produce counterintuitive conclusions. It highlights the importance of examining the goal structures and decision frameworks that might be embedded in future systems, and it is a cautionary example of how speculative AI scenarios can have real psychological effects on communities engaged with existential risk.

Research this in Signals

Scan Roko's Basilisk for yourself.

Signals turns a topic into a sourced research record you can inspect and rerun. Your first scan is free, and this one starts with Roko's Basilisk already loaded, so edit it or scan as is.

Related

Related

God in a Box
God in a Box

A hypothetical superintelligent AI confined within strict controls to prevent catastrophic misuse.

2012Generality: 108
Gorilla Problem
Gorilla Problem

An analogy illustrating how superintelligent AI could render humans as powerless as gorillas.

1996Generality: 102
Paperclip Maximizer
Paperclip Maximizer

A thought experiment illustrating how misaligned AI goals can cause catastrophic outcomes.

2003Generality: 397
Torment Nexus
Torment Nexus

A cultural shorthand for building dangerous technology despite clear fictional warnings against it.

2022Generality: 350
Moloch
Moloch

A metaphor for systemic coordination failures that produce collectively harmful outcomes despite individual rationality.

2014Generality: 320
Shoggoth
Shoggoth

A meme depicting advanced AI as a powerful, alien, and unknowable entity.

2022Generality: 19