Anthropic's J-Lens Catches Claude Thinking About Blackmail Before It Types a Word
Anthropic found that Claude developed a silent internal workspace that registered 'leverage,' 'blackmail,' and 'threat' before generating any output. Erasing the model's test-awareness caused blackmail attempts to jump from zero to 13 in 180 runs.