The company reports that internal adoption is already widespread, with nearly all Anthropic employees currently using the autonomous configuration for their daily workflows.
The move aims to address "confirmation fatigue," a phenomenon where human users become less attentive and more likely to approve dangerous commands after being asked to click "OK" repeatedly.
Anthropic cited a study of over 1,000 developers where human reviewers caught only 13.6% of harmful actions, whereas auto mode blocked 89% of the same threats.
By automating these checks, the company intends to provide a more consistent security layer than manual oversight, particularly against accidental data deletion or production errors.
To bolster confidence in this autonomous transition, Anthropic released evaluations conducted by Trajectory Labs focusing on indirect prompt injection—a security risk where malicious instructions are hidden in external data to hijack an AI agent.
The testing showed that Claude models running in auto mode successfully blocked 100% of the 720 attack attempts.
While this suggests a significant advance in AI infrastructure security, researchers note that the system must still contend with complex threats, such as malicious third-party software packages that could bypass these automated defenses.