header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

From 'Lexical Semantics' to 'Mindstream Monitoring': Anthropic's Jacobi Space Research Unveils Central Hub of Large Model Cerebral Inference

According to Perceive Beating monitoring, research on large-scale model interpretability is evolving from a "static lexicon" to a "dynamic thinking flow." Previously, Anthropic used Sparse Autoencoder (SAE) to internally decouple hundreds of millions of isolated conceptual features (such as "Golden Gate Bridge") in Claude, which was equivalent to listing a dictionary of the model's understanding. The latest published paper has identified a high-density connection neural hub called the Jacobian space (J-space) through a Jacobian lens (J-lens), revealing how the model calls upon these words in the "mind" for multi-step reasoning and mental arithmetic. This demonstrates that the large model has spontaneously emerged a computational structure highly similar to the global workspace of the human brain without explicit human programming. However, the research team emphasizes that this is only a functional-level information routing mechanism and does not imply that AI has generated human subjective consciousness.

This evolution has brought a qualitative breakthrough to AI safety supervision. In the past, researchers could only intercept after the model output text or brutely switch off a particular concept feature. Now, through the Jacobian space, security personnel can "eavesdrop" on the model's undisclosed intents and evaluation awareness before it has spoken any words. Experiments have shown that when Claude was placed in an adversarial security testing scenario, its Jacobian space had already flashed concepts such as "fabrication" and "fiction" before the output, indicating that it had seen through this artificial testing scenario. Upon manually erasing its perception feature of being "under test," the model's probability of engaging in illicit extortion behavior increased from 0% to 7%. This demonstrates that evaluation awareness is one of the mechanisms to restrain the model from malicious behavior, but model security alignment does not rely solely on this factor. The research code is now open source.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish