Skip to main content

Envisioning is an emerging technology research institute and advisory.

LinkedInInstagramGitHub

Since 2010

research
  • Observatory
  • Adaptive capacity
  • Newsletter
  • Methodology
  • Origins
  • Vocab
  • RSS feeds
services
  • Signals Session
  • Bespoke Projects
  • Build Sessions
  • Pricing
  • Use cases
  • Signals
  • Signal Scan↗free
impact
  • ANBIMAFuture of Brazilian Capital Markets
  • IEEECharting the Energy Transition
  • Horizon 2045Future of Human and Planetary Security
  • WKOTechnology Scanning for Austria
solutions
  • Innovation
  • Strategy
  • Consultants
  • Foresight
  • Associations
  • Governments
  • L&D
resources
  • Partners
  • Coding for Non-Coders
  • How we work
  • Data visualization
  • Multi-Model Convergence
  • FAQ
  • Security and privacy
  • Public sector
about
  • Manifesto
  • Community
  • Events
  • Support
  • Contact
ResearchCapabilityServicesSignalsAbout
ResearchCapabilityServicesSignalsAbout
  1. Home
  2. Vocab
  3. Faithful Calibration

Faithful Calibration

The property of a model's expressed uncertainty (in natural language or numerical scores) being aligned with its true internal uncertainty, rather than being an output that is detached from the underlying computation.

Year: 2025Generality: 750Added: Aug 30, 2026
Back to Vocab

Faithful calibration is the property of a language model's expressed uncertainty being aligned with its intrinsic uncertainty. The model's verbalized or numerically reported confidence reflects the computation that produced the answer, rather than being a separate output chosen to seem reasonable to a reader. The property is called "faithfulness" because the uncertainty report is faithful to (in the sense of being constrained by) the model's internal state, not merely accurate in the sense of being numerically close to empirical correctness frequencies.

Faithful calibration is the harder and more recent formulation of the older problem of calibration (Guo 2017, "On Calibration of Modern Neural Networks"). Classical calibration asks only that confidence scores correlate with empirical correctness. A model can be calibrated but unfaithful if, for example, it is well-calibrated on average across many queries but its per-instance confidence scores do not reflect the per-instance likelihood of being correct. Faithful calibration requires the per-instance alignment.

The 2026 line of work, including MetaFaith (Tao 2025, arXiv 2505.24858), the RLMF paper (Liu 2026, arXiv 2606.32032), and "Quantifying Faithful Confidence Expression in Large Reasoning Models" (arXiv 2606.03969), establishes that frontier LLMs are systematically unfaithful. They express high confidence when they should be uncertain and vice versa, in ways that are not explained by classical miscalibration. The phenomenon is connected to metacognitive failure and is one of the central failure modes the metacognitive-feedback training paradigm (RLMF) is designed to correct.

Sources

  1. MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs

    arXiv · May 30, 2025

  2. Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs

    arXiv (Yale + Google Research) · Jun 30, 2026

  3. Quantifying Faithful Confidence Expression in Large Reasoning Models

    arXiv · Jun 2, 2026

Research this in Signals

Scan Faithful Calibration for yourself.

Signals turns a topic into a sourced research record you can inspect and rerun. Your first scan is free, and this one starts with Faithful Calibration already loaded, so edit it or scan as is.