Skip to main content

Envisioning is a research institute that studies how institutions adapt to technological change.

LinkedInInstagramGitHub

Since 2010

research
  • Observatory
  • Adaptive capacity
  • Newsletter
  • Methodology
  • Origins
  • Vocab
  • RSS feeds
services
  • Signals Session
  • Bespoke Projects
  • Build Sessions
  • Pricing
  • Use cases
  • Signals
  • Signal Scan↗free
impact
  • ANBIMAFuture of Brazilian Capital Markets
  • IEEECharting the Energy Transition
  • Horizon 2045Future of Human and Planetary Security
  • WKOTechnology Scanning for Austria
solutions
  • Innovation
  • Strategy
  • Consultants
  • Foresight
  • Associations
  • Governments
  • L&D
resources
  • Partners
  • Coding for Non-Coders
  • How we work
  • Data visualization
  • Multi-Model Convergence
  • FAQ
  • Security and privacy
  • Public sector
about
  • Manifesto
  • Community
  • Events
  • Support
  • Contact
ResearchCapabilityServicesSignalsAbout
ResearchCapabilityServicesSignalsAbout
  1. Home
  2. Vocab
  3. RLCD (Reinforcement Learning for Calibrated Decisions)

RLCD (Reinforcement Learning for Calibrated Decisions)

A reinforcement-learning based training method, introduced by TypeSafe AI, that optimizes a model's predicted probabilities to match real-world outcome frequencies rather than optimizing for human preference.

Year: 2026Generality: 250Added: Sep 18, 2026
Back to Vocab

Reinforcement Learning for Calibrated Decisions (RLCD) is a training method TypeSafe AI says it used to build Jev, its first "System One" model. The company announced Jev in a September 2026 blog post by cofounder and CEO Diogo Almeida. Reinforcement learning from human feedback (RLHF), a method Almeida co-developed at OpenAI for InstructGPT, trains a model to produce outputs that human raters prefer. RLCD instead optimizes a model's output probabilities so that they are calibrated: among all the predictions the model assigns a given confidence level, such as 90%, that fraction of them should actually turn out correct.

Calibration in this sense says nothing about whether any single prediction is right. It only means the model's stated confidence tracks its accuracy in aggregate across many predictions. TypeSafe treats this as distinct from a model simply being correct or incorrect. Because Jev's outputs are constrained to a fixed schema, a malformed or out-of-schema answer counts as a separate failure mode from a well-formed but wrong one. RLCD targets the confidence attached to well-formed answers specifically, aiming to make it trustworthy enough that software can act on it automatically above a chosen threshold.

As of its announcement, RLCD has been described only in TypeSafe's own blog post and subsequent press coverage. No peer-reviewed paper or technical report documents its exact mechanism, so the loss function, reward model, and calibration procedure remain undisclosed. Commentators have noted this distinction, describing the architecture, sampler, and training method as company claims rather than verified technical results pending a fuller technical writeup.

Sources

  1. Introducing System One Models & Jev

    TypeSafe AI · Sep 15, 2026

  2. TypeSafe AI debuts model for machines that plays Doom

    The Register · Sep 16, 2026

  3. A deep dive into Jev, TypeSafe's System One model

    flaviocopes.com · Sep 17, 2026

Research this in Signals

Scan RLCD (Reinforcement Learning for Calibrated Decisions) for yourself.

Signals turns a topic into a sourced research record you can inspect and rerun. Your first scan is free, and this one starts with RLCD (Reinforcement Learning for Calibrated Decisions) already loaded, so edit it or scan as is.