Demystifying Anthropic's J-Space: A Mathematical Primer

| Source: Towards Data Science

Tags: Claude Opus 4.6, Anthropic, interpretability, mechanistic interpretability, J-space, Global Workspace Theory, alignment

Anthropic's J-space — demonstrated on Claude Opus 4.6 — is a geometric subspace of the residual stream capturing "verbalizable" representations, analogous to the human brain's global workspace. This mathematical primer unpacks why J-space is a union of k-sparse non-negative cones (not a linear subspace), and how it functions as a principled alignment auditing tool for LLMs.

Details

Anthropic's paper "Verbalizable Representations Form a Global Workspace in Language Models" draws a formal parallel between Claude Opus 4.6's internal representations and Global Workspace Theory (GWT) — the neuroscience framework hypothesized to underpin human consciousness. The J-space is their proposed name for this workspace inside the model. This Towards Data Science primer fills in the math the original paper leaves implicit. J-lens vectors are derived by measuring each token representation's average causal impact on output logits, then projecting through the unembedding matrix. There is one J-lens vector per vocabulary token per layer. The J-space is formally defined as the union of k-sparse non-negative cones over these vectors — a non-linear, non-subspace geometric object. In a 4096-dimensional residual stream, pairs of J-lens vectors are nearly orthogonal (pairwise inner product ~0.015), with the exception of semantically related tokens such as "king" and "emperor." This concentration-of-measure property from high-dimensional geometry is what makes the J-space well-behaved despite being an overcomplete frame. Representations that fall inside J-space are those the model can articulate; those outside are non-verbalizable internal computations. For mechanistic interpretability practitioners, this provides a structured framework to audit whether model reasoning is accessible and to target surgical interventions — directly relevant to alignment research on Claude-family models.