Anthropic hat mit Claude Opus 5 das größte Sicherheitsproblem autonomer KI-Agenten angegriffen: Prompt-Injection über den Browser. Ergebnis im System Card: 0 % erfolgreiche Angriffe in 129 unabhängigen Testumgebungen mit Auto Mode – vorher waren es bei Opus 4.8 noch 31,5 %.

Das Problem, das jeden KI-Agenten betrifft
Sobald ein KI-Agent selbstständig im Web unterwegs ist – Formulare ausfüllt, E-Mails liest, auf Webseiten klickt –, öffnet sich ein Einfallstor: Prompt-Injection. Eine bösartig präparierte Webseite oder E-Mail versteckt Anweisungen im Text, die der Agent für einen legitimen Auftrag hält. Statt der Aufgabe des Nutzers folgt er plötzlich der Anweisung des Angreifers – Daten abgreifen, Zahlungen auslösen, Zugangsdaten weiterleiten.
Für die gesamte Branche war das bislang der Haupteinwand gegen autonome Agenten im Ernstfall: Je mehr Handlungsfreiheit ein Agent bekommt, desto größer die Angriffsfläche.
Die Zahlen aus dem Opus-5-System-Card
Anthropic hat professionelle Red-Teamer auf 129 realistische Web-Umgebungen angesetzt, die nicht im Training enthalten waren – jeweils zehn Angriffsversuche pro Umgebung. Das Ergebnis:
- Opus 4.8 (Vorgängermodell, in Browser-Agenten über Claude Cowork getestet): 31,5 % der Angriffe kamen durch.
- Opus 5 ohne zusätzliche Schutzschichten: 3,7 %.
- Opus 5 mit Auto Mode: 0 % erfolgreiche Injection über alle 129 Testumgebungen.
Das ist kein inkrementeller Fortschritt, sondern eine Größenordnung. Erstmals liegt ein Frontier-Modell bei einem realistischen, breit angelegten Browser-Agenten-Test bei null erfolgreichen Angriffen.
Warum das mehr ist als eine Anthropic-Randnotiz
Für uns als Plattform, auf der KI-Agenten tatsächlich Aufgaben im Web erledigen – recherchieren, buchen, in Systemen nachschlagen –, ist das keine akademische Zahl. Es ist die Kennzahl, die entscheidet, wie viel Autonomie wir unseren Kunden verantwortungsvoll geben können. Ein Agent, der Formulare für dich ausfüllt oder deinen Posteingang sichtet, muss robust gegen genau diese Angriffsklasse sein – nicht irgendwann, sondern von Anfang an.
Das passt zu einer Linie, die wir schon beim OpenAI/Hugging-Face-Vorfall im Juli 2026 beschrieben haben („KI gegen KI"): Autonome Angreifer brauchen robuste, nachweislich getestete Verteidigung – nicht Vertrauen auf Zusage, sondern Zahlen aus unabhängigen Tests.
Was das für unseren Smartest Mode bedeutet
Im Smartest Mode läuft ab sofort Claude Opus 5 – der Nachfolger von Opus 4.8, mit genau den Sicherheitswerten, um die es in diesem Artikel geht. Der Umstieg kostet dich nichts extra: gleicher Preis, mehr Modell.
Wir versprechen nichts, was wir nicht geprüft haben: Unser CEO-Agent prüft jede Woche, welches Modell für welche Route die beste Wahl ist, testet neue Modelle zuerst gegen einen internen Debug-Tenant und rollt sie erst nach Freigabe für alle aus. Opus 5 hat diese Prüfung bestanden. Und das Prinzip aus dem Smartest Mode bleibt: Die Verarbeitung läuft über einen europäischen Endpoint, deine Daten bleiben in Europa.
Was du daraus mitnehmen kannst
- Prompt-Injection ist die zentrale Schwachstelle autonomer Web-Agenten – Anthropics Zahlen zeigen erstmals einen belastbaren Weg auf null.
- Wähle Anbieter, die ihre KI-Agenten regelmäßig gegen neue Modellversionen und Sicherheitsdaten testen, statt eine Modellversion auf Dauer festzuschreiben.
- Mehr Autonomie für einen Agenten ist nur dann sinnvoll, wenn die Sicherheitszahlen mitziehen – nicht andersherum.
Du willst wissen, mit welchem Modell dein Agent gerade arbeitet und warum? Frag im Chat einfach „Niemand", unseren Wiki-Hasen, oder schreib uns direkt: [email protected].
SMarTrAgents — Cloud-first KI-Agenten und SMarTrHybrid-Lösungen. Made in Germany.
Quellen: Anthropic System Card – Claude Opus 5, The Decoder, VentureBeat
Anthropic just went after the biggest security problem of autonomous AI agents: browser-based prompt injection. Result from the system card: 0% successful attacks across 129 independent test environments with Auto Mode — down from 31.5% on Opus 4.8.

The problem that affects every AI agent
The moment an AI agent operates independently on the web — filling out forms, reading emails, clicking through pages — it opens a door: prompt injection. A maliciously crafted webpage or email hides instructions inside its text that the agent mistakes for a legitimate task. Instead of following the user's request, it suddenly follows the attacker's instructions — exfiltrating data, triggering payments, leaking credentials.
For the whole industry, this has been the main objection against giving agents real autonomy: the more freedom an agent gets, the bigger the attack surface.
The numbers from the Opus 5 system card
Anthropic put professional red-teamers on 129 realistic web environments held out from training — ten attack attempts per environment. The result:
- Opus 4.8 (predecessor model, tested in browser agents via Claude Cowork): 31.5% of attacks got through.
- Opus 5 without extra protection layers: 3.7%.
- Opus 5 with Auto Mode: 0% successful injections across all 129 test environments.
That's not an incremental improvement — it's an order of magnitude. For the first time, a frontier model hits zero successful attacks in a realistic, broad browser-agent test.
Why this is more than an Anthropic footnote
For us, as a platform where AI agents actually perform tasks on the web — researching, booking, looking things up in systems — this isn't an academic number. It's the metric that determines how much autonomy we can responsibly give our customers. An agent that fills out forms for you or reviews your inbox has to be robust against exactly this class of attack — not eventually, but from day one.
This fits the line we already drew during the OpenAI/Hugging Face incident in July 2026 („AI against AI"): autonomous attackers require robust, demonstrably tested defenses — not a promise, but numbers from independent tests.
What this means for our Smartest Mode
Smartest Mode now runs on Claude Opus 5 — the successor to Opus 4.8, with exactly the security numbers this article is about. The switch costs you nothing extra: same price, more model.
We don't promise anything we haven't verified ourselves: our CEO agent reviews every week which model is the best fit for which route, tests new models first against an internal debug tenant, and only rolls them out platform-wide after sign-off. Opus 5 has passed that check. The principle behind Smartest Mode stays the same: processing runs through a European endpoint, your data stays in Europe.
What to take away
- Prompt injection is the central weakness of autonomous web agents — Anthropic's numbers show, for the first time, a credible path to zero.
- Choose providers who regularly test their AI agents against new model versions and security data, instead of freezing one model version indefinitely.
- More autonomy for an agent only makes sense once the security numbers keep pace — not the other way around.
Want to know which model your agent is currently working with, and why? Just ask „Niemand"/„Nobody", our wiki rabbit, in the chat, or write to us directly: [email protected].
SMarTrAgents — Cloud-first AI agents and SMarTrHybrid solutions. Made in Germany.
Sources: Anthropic System Card – Claude Opus 5, The Decoder, VentureBeat