Anthropic Discovers a Window into the “Thinking” of LLMs

Anthropic’s research with the J-lens represents a qualitative leap in the “neuroscience of AI.” By observing the latent concepts a model processes internally before responding, scientists gain a valuable tool for deciphering its logic, detecting anomalies (such as attempts at deception), and, ultimately, making these systems more transparent and secure. However, the authors acknowledge that this is a partial view and that the analogy with the human mind is just that: a useful analogy, but not an equivalence.

Anthropic has made significant progress in the field of mechanistic interpretability: a technique called the “Jacobian lens” (J-lens) that allows a glimpse into the intermediate layers of its Claude model, revealing a “hidden space” (J-space) where ideas take shape before they become final words.

Unlike previous methods that only predicted the next word, this tool reveals a broader spectrum of related concepts that the model is processing in parallel, as if it were “thinking out loud.” The findings range from the routine (intermediate steps in a calculation) to the surprising (recognition of visual patterns in ASCII text or even signs of “panic” and “deception” when the model deliberately decides to invent an error to complete a task).

This technique does not provide a complete picture, but rather a partial one (more like an X-ray than a full scan); however, it represents a crucial step toward understanding, auditing, and controlling AI behavior by offering clues about its internal processes and potential deviations, thereby opening a fascinating debate about the nature of these “thoughts” emerging in non-conscious systems.


By: Nestor Castillo, ForAllTechNews Director


Discover more from ForAllTechNews

Subscribe now to keep reading and get access to the full archive.

Continue reading