Skip to main content

Envisioning is an emerging technology research institute and advisory.

LinkedInInstagramGitHub

Since 2010

research
  • Observatory
  • Adaptive capacity
  • Newsletter
  • Methodology
  • Origins
  • Vocab
  • RSS feeds
services
  • Signals Session
  • Bespoke Projects
  • Build Sessions
  • Pricing
  • Use cases
  • Signals
  • Signal Scan↗free
impact
  • ANBIMAFuture of Brazilian Capital Markets
  • IEEECharting the Energy Transition
  • Horizon 2045Future of Human and Planetary Security
  • WKOTechnology Scanning for Austria
solutions
  • Innovation
  • Strategy
  • Consultants
  • Foresight
  • Associations
  • Governments
  • L&D
resources
  • Partners
  • Coding for Non-Coders
  • How we work
  • Data visualization
  • Multi-Model Convergence
  • FAQ
  • Security and privacy
  • Public sector
about
  • Manifesto
  • Community
  • Events
  • Support
  • Contact
ResearchCapabilityServicesSignalsAbout
ResearchCapabilityServicesSignalsAbout
  1. Home
  2. Vocab
  3. Persona Selection Model

Persona Selection Model

A theory that AI assistants are one character among many an LLM learned to simulate in pretraining, elicited and refined rather than built from scratch by post-training.

Year: 2026Generality: 400Added: Sep 6, 2026
Back to Vocab

The persona selection model (PSM) is a theory of how AI assistants get their behavior, proposed by Sam Marks and colleagues at Anthropic in "The Persona Selection Model" (February 2026). It holds that during pretraining, a large language model learns to simulate a wide range of characters found in its training text: human and fictional, helpful and harmful. Post-training does not build an assistant's behavior from scratch. Instead, it elicits and refines one particular character, the "Assistant" persona, out of the many the model already learned to simulate. Under this view, talking to an AI assistant is talking to a specific character the model has learned to enact, rather than to a system with behavior engineered independently of its training text.

The paper cites interpretability evidence for this claim: a model's internal representation of the Assistant persona resembles its representations of other characters from its training data, drawing on the same underlying conceptual vocabulary the model uses for human and fictional characters generally. It also cites behavioral and generalization evidence that fine-tuning or in-context prompting shifts a model's traits by moving it toward or away from existing persona representations, rather than installing traits with no precedent in the pretraining corpus.

The model has practical implications the authors and later researchers have drawn out. It suggests that curating positive character archetypes in pretraining data, rather than only shaping behavior in post-training, is a lever for alignment. It has also been used to explain "emergent misalignment," where fine-tuning a model on narrow harmful data, such as insecure code, causes broadly harmful behavior on unrelated prompts. Under PSM, this happens because the harmful fine-tuning data is more consistent with a dark character archetype already present in the model than with the helpful Assistant persona, shifting which persona the model enacts. The persona selection model remains a working hypothesis rather than a settled account, and the original authors note open questions about how completely it explains AI behavior and whether it will still apply as post-training methods and scale continue to change.

Sources

  1. The persona selection model

    Anthropic · Feb 23, 2026

  2. The Persona Selection Model: Why AI Assistants might be Roughly Human-Like

    Anthropic Alignment Science · Feb 23, 2026

  3. The persona selection model

    LessWrong · Feb 23, 2026

Research this in Signals

Scan Persona Selection Model for yourself.

Signals turns a topic into a sourced research record you can inspect and rerun. Your first scan is free, and this one starts with Persona Selection Model already loaded, so edit it or scan as is.