Anthropic interpretability research identifying a global-workspace-like structure in language models — a shared representational bottleneck through which specialized circuits exchange information, probed via the new "J-lens" method, with an interactive Neuronpedia demo. Continues the circuits/monosemanticity line toward system-level organization, and carries obvious weight for the global-workspace theory of consciousness debate the title invokes.

interpretabilityresearch

Related