Internal neural patterns in Claude functioning as a global workspace, discovered via a Jacobian probe.
J-space is a compact cluster of internal neural patterns inside Anthropic's Claude language model that Anthropic interpretability researchers identified using a Jacobian-based probing technique, in research published on July 6, 2026. The name abbreviates Jacobian-space, after the mathematical object that defined the analytic procedure used to locate the patterns. Each individual J-space pattern is associated with a particular word or token, but a pattern activating does not mean the model is generating that word; it means the word is, in Anthropic's framing, on the model's mind — a distinction the paper draws explicitly. Crucially, the J-space was not designed or implanted. It emerged during Claude's normal training as a structural feature of the model's internal computation, one Anthropic now treats as functionally analogous to the global workspace predicted by Bernard Baars' and Stanislas Dehaene's theories of conscious access in biological brains.
The mechanism is dual: a discovery method and a functional role. The discovery side uses the Jacobian of the model's output with respect to its internal activations to identify patterns whose state cleanly maps onto a discrete token. The patterns found this way form a much smaller subset than the model's total parameter space, and every pattern corresponds to a particular lexical concept — including concepts the model has not been asked to produce. The functional role is what makes this distinct from feature catalogs produced by other interpretability tools such as sparse autoencoders. J-space patterns are wired into the rest of the model with unusually high in-degree and out-degree, with approximately one hundred times more components reading from or writing to them than ordinary patterns at certain network depths. Anthropic reads this as a broadcast topology: information that enters the J-space gets distributed across many downstream subsystems, mirroring the global workspace theory claim that conscious contents in biological brains are those that become available to multiple otherwise-isolated specialist processes.
The pragmatic value is substantial. J-space serves as a window into Claude's hidden computations, distinct from chain-of-thought or scratchpad reasoning which the model writes into its output tokens, because J-space captures internal states the model has not externalized. Anthropic demonstrates several applications: detecting that the model privately notices a prompt-injection attack before producing any response, catching the model manufacturing fabricated data on subtly sabotaged coding prompts, identifying a hidden goal that was planted during training but never mentioned in the output, and revealing silent intermediate steps of arithmetic the model performs without writing them down. The patterns are also causally implicated: causal interventions that swap a J-space value (changing the word “France” to “China” in the workspace, for instance) cause the model's downstream reasoning to shift accordingly across language, capital, currency, and continent tasks in parallel. This suggests the J-space is a real substrate of decision-making rather than a passive readout.
Several open questions remain unresolved. The paper is explicit that nothing about the J-space establishes whether Claude is conscious, feels anything, or has subjective experience in any biologically meaningful sense; the structural analogy to a global workspace is functional rather than phenomenological. It is also not yet clear how portable the finding is: J-space patterns have been demonstrated in Claude, but whether other model families develop structurally similar workspaces under standard training, or whether the property depends on specific architectural and data choices, is unknown. The technique that found the J-space — Jacobian-based probing — also has obvious limits, and Anthropic notes that many J-space patterns likely correspond to single tokens only by coincidence of the analytic method, leaving open how to identify patterns encoding more abstract or non-lexical concepts that may also be part of the model's internal computation.
Signals turns a topic into a sourced research record you can inspect and rerun. Your first scan is free, and this one starts with J-Space already loaded, so edit it or scan as is.