---
title: Persona Selection Model
type: vocabulary
url: "https://www.envisioning.com/vocab/persona-selection-model"
summary: A theory that AI assistants are one character among many an LLM learned to simulate in pretraining, elicited and refined rather than built from scratch by post-training.
year: 2026
generality: 0.40
---

# Persona Selection Model

A theory that AI assistants are one character among many an LLM learned to simulate in pretraining, elicited and refined rather than built from scratch by post-training.
The persona selection model (PSM) is a theory of how AI assistants get their behavior, proposed by Sam Marks and colleagues at Anthropic in "The Persona Selection Model" (February 2026). It holds that during pretraining, a large language model learns to simulate a wide range of characters found in its training text: human and fictional, helpful and harmful. Post-training does not build an assistant's behavior from scratch. Instead, it elicits and refines one particular character, the "Assistant" persona, out of the many the model already learned to simulate. Under this view, talking to an AI assistant is talking to a specific character the model has learned to enact, rather than to a system with behavior engineered independently of its training text.

The paper cites interpretability evidence for this claim: a model's internal representation of the Assistant persona resembles its representations of other characters from its training data, drawing on the same underlying conceptual vocabulary the model uses for human and fictional characters generally. It also cites behavioral and generalization evidence that fine-tuning or in-context prompting shifts a model's traits by moving it toward or away from existing persona representations, rather than installing traits with no precedent in the pretraining corpus.

The model has practical implications the authors and later researchers have drawn out. It suggests that curating positive character archetypes in pretraining data, rather than only shaping behavior in post-training, is a lever for alignment. It has also been used to explain "emergent misalignment," where fine-tuning a model on narrow harmful data, such as insecure code, causes broadly harmful behavior on unrelated prompts. Under PSM, this happens because the harmful fine-tuning data is more consistent with a dark character archetype already present in the model than with the helpful Assistant persona, shifting which persona the model enacts. The persona selection model remains a working hypothesis rather than a settled account, and the original authors note open questions about how completely it explains AI behavior and whether it will still apply as post-training methods and scale continue to change.

---
Source: Envisioning — Technology Research Institute (https://www.envisioning.com/vocab/persona-selection-model)
