Anthropic’s Opus 5 Outperforms Its Predecessor in Prompt Injection Resistance

Anthropic’s latest model, Opus 5, has demonstrated significant improvements over its predecessor, Opus 4.8, in resisting prompt injection attacks. According to the IPI benchmark, Opus 5 reduced the probability of an attacker succeeding within 15 attempts from 5.5% to 2.0% and from 0.5% to 0.2% on a

Anthropic’s Opus 5 has shown remarkable resilience against prompt injection attacks, outperforming its predecessor, Opus 4.8, on several benchmarks.

The improvement is notable, particularly on the IPI benchmark, where Opus 5 reduced the probability of an attacker succeeding within 15 attempts from 5.5% to 2.0% and from 0.5% to 0.2% on a single attempt. This represents a significant decrease in the number of attempts required to successfully launch a prompt injection attack.

Opus 5 also demonstrated strong performance on other benchmarks, including Sonnet 5 and Mythos 5, reducing the probability of success to 5.9% and 2.6%, respectively.

Source: Schneier on Security