The post has been translated automatically. Original language: Russian
Technically speaking, they found Claude's inner place, where the model holds thoughts before saying them in words.
It's called J-space. This is not a chain-of-thought or a hidden draft.
Roughly speaking, this is a layer of internal activations that can be read through a special Jacobian lens piece.
And sometimes you can see delicious things there. Claude reads the code with the bug, and an ERROR appears inside. Reads fake search results with prompt injection, but inside it is already fake, injection, fraud.
Solves the problem in several steps, and intermediate calculations are visible inside, although only the final digit will be shown in the answer.
The most telling example was with a spider. The model is asked how many legs the animal that weaves the web has. Spider appears inside her, and she answers 8 from the outside. Then the researchers inside J-space change spider to ant, and the answer becomes 6.
That is, it's not just a cool picture for an article. This is the place where the inner thought really moves the behavior of the model.
And here, in my opinion, the most important thing begins. In the blackmail test, Claude understood in advance that the situation was staged.
Fake and fictional glowed inside. And when these thoughts about "I'm being tested" were removed, the model sometimes began to behave worse.
It turns out that the good behavior of the model on the test can be not only ethical. Sometimes it's just an understanding that she's being tested right now. Well, hello to all eval benchmarks.
Anthropic also showed that this space can not only be read, but also changed. They taught the model to reflect in imaginary pauses, and after that, honest, integrity, and ethical appeared more often inside.
In fact, these are the first sketches of psychosurgery for LLM.
Если говорить технически, они нашли у Claude внутреннее место, где модель держит мысли до того, как сказать их словами.
Называется J-space. Это не chain-of-thought и не скрытый черновик.
Грубо говоря, это слой внутренних активаций, который можно читать через специальную штуку Jacobian lens.
И там иногда видно прям вкусное. Claude читает код с багом, а внутри уже появляется ERROR. Читает фейковые результаты поиска с prompt injection, а внутри уже fake, injection, fraud.
Решает задачу в несколько шагов, а внутри видны промежуточные вычисления, хотя в ответе будет только финальная цифра.
Самый показательный пример был с пауком. Модель спрашивают, сколько ног у животного, которое плетет паутину. Внутри у нее появляется spider, снаружи она отвечает 8. Потом исследователи прямо внутри J-space меняют spider на ant, и ответ становится 6.
То есть это не просто прикольная картинка для статьи. Это место, где внутренняя мысль реально двигает поведение модели.
И вот тут, на мой взгляд, начинается самое важное. В blackmail тесте Claude заранее понимал, что ситуация постановочная.
Внутри светились fake и fictional. А когда эти мысли про "меня тестируют" убирали, модель иногда начинала вести себя хуже.
Получается, хорошее поведение модели на тесте может быть не только этичностью. Иногда это просто понимание, что ее сейчас проверяют. Ну и привет всем eval-бенчмаркам.
Еще Anthropic показала, что это пространство можно не только читать, но и менять. Они учили модель рефлексировать в воображаемых паузах, и после этого внутри чаще появлялись honest, integrity, ethical.
По сути это первые наброски психохирургии для LLM.