Interpretability: a "global workspace" spotted in Claude's internal representations
Anthropic uses its Jacobian lens to find a global workspace in Claude, revealing internal thoughts the AI is ready to verbalize but does not write out.
Anthropic has published an interpretability study introducing the Jacobian lens (J-lens), a technique that identifies internal representations a model is ready to verbalize without necessarily writing them out. This set of directions, dubbed J-space, would, according to the authors, form a "global workspace" comparable to that described by a neuroscientific theory of conscious access. This subspace accounts for less than 10% of the internal activity at each layer, but researchers observe that it concentrates what the model can report, manipulate, and mobilize for reasoning, while the rest of the processing would remain automatic. Anthropic specifies that this structure was not programmed: it emerges spontaneously during training.
The team primarily sees it as an alignment auditing tool. The J-lens would allow reading "thoughts" that the model does not express: identifying a prompt injection, fabricated data, or the fact that it has recognized a situation as a test. The authors emphasize that the method remains imperfect and captures only a fraction of the reasoning.
The comparison with theories of consciousness is framed: Anthropic distinguishes conscious access, which is purely functional, from phenomenal consciousness, on which the study does not comment. External commentaries accompany the publication, including those from neuroscientists Stanislas Dehaene and Lionel Naccache, originators of the global workspace theory, researchers from Eleos AI, and Neel Nanda (Google DeepMind), who reports a partial replication of the results on an open-weights model.