Opus 5: From 31.5 Percent to 0 Percent Successful Prompt Injection

Anthropic just went after the biggest security problem of autonomous AI agents: browser-based prompt injection. Result from the system card: 0% successful attacks across 129 independent test environments with Auto Mode, down from 31.5% on Opus 4.8.
The problem that affects every AI agent
The moment an AI agent operates independently on the web, filling out forms, reading emails, clicking through pages, it opens a door: prompt injection. A maliciously crafted webpage or email hides instructions inside its text that the agent mistakes for a legitimate task. Instead of following the user’s request, it suddenly follows the attacker’s instructions, exfiltrating data, triggering payments, leaking credentials.
For the whole industry, this has been the main objection against giving agents real autonomy: the more freedom an agent gets, the bigger the attack surface.
The numbers from the Opus 5 system card
Anthropic put professional red-teamers on 129 realistic web environments held out from training, ten attack attempts per environment. The result:
- Opus 4.8 (predecessor model, tested in browser agents via Claude Cowork): 31.5% of attacks got through.
- Opus 5 without extra protection layers: 3.7%.
- Opus 5 with Auto Mode: 0% successful injections across all 129 test environments.
That’s not an incremental improvement, it’s an order of magnitude. For the first time, a frontier model hits zero successful attacks in a realistic, broad browser-agent test.
Why this is more than an Anthropic footnote
For us, as a platform where AI agents actually perform tasks on the web, researching, booking, looking things up in systems, this isn’t an academic number. It’s the metric that determines how much autonomy we can responsibly give our customers. An agent that fills out forms for you or reviews your inbox has to be robust against exactly this class of attack, not eventually, but from day one.
This fits the line we already drew during the OpenAI/Hugging Face incident in July 2026 („AI against AI”): autonomous attackers require robust, demonstrably tested defenses, not a promise, but numbers from independent tests.
What this means for our Smartest Mode
Smartest Mode now runs on Claude Opus 5: the successor to Opus 4.8, with exactly the security numbers this article is about. The switch costs you nothing extra: same price, more model.
We don’t promise anything we haven’t verified ourselves: our CEO agent reviews every week which model is the best fit for which route, tests new models first against an internal debug tenant, and only rolls them out platform-wide after sign-off. Opus 5 has passed that check. And we keep saying openly where the computing happens: your account, your conversations and your files sit on servers inside the European Union, while the model execution in Smartest Mode currently runs at Anthropic in the USA. What that means for your data is set out in full in Smartest Mode: Opus 5 power for your AI agents.
What to take away
- Prompt injection is the central weakness of autonomous web agents, Anthropic’s numbers show, for the first time, a credible path to zero.
- Choose providers who regularly test their AI agents against new model versions and security data, instead of freezing one model version indefinitely.
- More autonomy for an agent only makes sense once the security numbers keep pace, not the other way around.
Want to know which model your agent is currently working with, and why? Just ask „Niemand”/„Nobody”, our wiki rabbit, in the chat, or write to us directly: [email protected].
SMarTrAgents, nine AI agents in one shared workspace. Accounts and files are held in the EU.
Sources: Anthropic System Card, Claude Opus 5, The Decoder, VentureBeat
Nine AI agents, one workspace
The agents draft quotes, keep the calendar and sort the inbox, and nothing goes out before you approve it.