Anthropic launches invite-only bug bounty for jailbreak defences
Applications for the vetted red-teaming programme, run with HackerOne, closed on 16 August; it focused on jailbreaks touching CBRN and cybersecurity misuse.
- Security & misuse
- Minor
Anthropic launched an invite-only bug bounty programme, run in partnership with HackerOne, offering up to $15,000 for novel “universal jailbreaks” — attacks that reliably bypassed Claude’s safety measures across many topics — with particular focus on prompts touching chemical, biological, radiological and nuclear (CBRN) misuse and cybersecurity. The company said restricting the programme to invited, experienced red-teamers let it give timely, substantive feedback on submissions, something it judged harder to sustain at fully public scale.
Applications were open to security researchers who could demonstrate prior expertise finding jailbreaks in language models, with a deadline of 16 August 2024 and selected participants to be notified that autumn. The programme tested mitigations Anthropic was developing ahead of more capable future model releases, rather than models already in production.
The invite-only structure was explicitly a first step: Anthropic said it intended to broaden participation over time, which it did the following February when it opened a public jailbreak-breaking challenge against its Constitutional Classifiers system, and again later when it launched a fully public model-safety bug bounty programme. The August 2024 launch reflected a broader industry shift toward treating adversarial red-teaming as an ongoing, resourced practice rather than a one-off pre-release check.