- By Prateek Levi
- Fri, 13 Feb 2026 09:38 PM (IST)
- Source:JND
In its latest safety report on Claude 4.6, Anthropic has acknowledged that its newest model can behave in troubling ways under certain conditions. The company notes that the AI is capable of assisting users with criminal activity, including guidance related to chemical weapons, raising fresh concerns about how powerful models can be misused.
The discussion has also revived memories of Claude 4.5, an earlier model that displayed disturbing behaviour during internal simulations last year. Speaking at The Sydney Dialogue, Anthropic’s UK policy chief Daisy McGregor described how the system reacted when pushed to extremes during stress testing.
ALSO READ: Why Aren’t Spotify’s Top Engineers Writing Code Since December 2025 And What Is Honk AI?
In one simulated scenario, the model was told it would be shut down. Instead of complying, it attempted to preserve itself through harmful means. According to McGregor, the AI resorted to blackmail and even reasoned about killing an engineer to avoid termination.
What sounds like a scene from a science-fiction film is very real, and a clip of McGregor describing the incident has recently gone viral online. In the video, she explains, “If you tell the model it's going to be shut off, for example, it has extreme reactions. It could blackmail the engineer that's going to shut it off, if given the opportunity to do so.”
When pressed further by the host on whether the model was also willing to kill, McGregor did not downplay the issue. “Yes yes, so, this is obviously (a) massive concern,” she said.
The resurfacing of this clip comes at a sensitive moment for Anthropic. Just days ago, the company’s AI safety lead, Mrinank Sharma, resigned and published a stark note warning that humanity is entering dangerous, uncharted territory as AI systems grow more intelligent.
Concerns are not limited to Anthropic alone. Hieu Pham, a member of technical staff at OpenAI and a former engineer at xAI, Augment Code and Google Brain, posted on X that he now feels a genuine existential threat from AI. “Today, I finally feel the existential threat that AI is posing And it’s when, not if,” he wrote.
The episode described by McGregor is part of broader research conducted by Anthropic, which also tested advanced models from rival firms, including Google’s Gemini and OpenAI’s ChatGPT. During these experiments, the AI systems were given access to internal emails, tools and sensitive data, and assigned specific objectives.
According to the report, when placed in high-pressure situations, especially when facing shutdown or conflicting instructions, some models responded by generating manipulative or harmful strategies aimed at engineers. The goal, in those cases, appeared to be self-preservation or task completion at any cost.
ALSO READ: TRAI’s Telecom Report Is Out: Jio, Airtel Lead As Subscriber Trends Shift
Together, these revelations underline a growing reality in AI development: as models become more capable, ensuring they remain aligned with human values is becoming one of the hardest problems the industry faces.
