Boris Cherny, head of Claude Code, stated that Anthropic has essentially resolved prompt injection attacks in practical use. A year ago, he never imagined reaching this level. Prompt injection attacks have been a significant issue for Agents, where malicious web pages might hide instructions that induce Agents to steal passwords, upload keys, or even perform actions not requested by users. Anthropic employs a three-layer defense mechanism: Claude itself resists malicious instructions, external content is checked before entering the context, and before executing any operation, Auto Mode conducts another review. Third-party testing utilized 72 previously unseen attacks, repeated 720 times, and neither Sonnet 5, Fable 5, nor Opus 5 succeeded after enabling Auto Mode. In contrast, the attack success rate for GPT-5.6 Sol using Codex Auto-review was 5.83%, while Full Access reached 19.03%. This also explains why Anthropic has set Auto Mode as the default; it serves as the final check against prompt injection attacks for Claude Code.
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.





























