Anthropic opens its bug bounty program to the public
Rewards run up to $10,000, and one track pays specifically for universal jailbreaks against Claude's Constitutional Classifiers rather than product bugs.
- Security & misuse
- Minor
Anthropic opened its HackerOne bug bounty programme to the public, ending an application-only, NDA-bound arrangement it had run privately with selected security researchers since 2024. Any registered HackerOne user can now submit reports.
The public programme runs two tracks. A Product Security track covers conventional vulnerabilities in Claude.ai, the Claude API and Claude Code, with rewards Anthropic said range up to $10,000 for critical findings. A separate Model Safety track pays specifically for novel, universal jailbreak attacks that defeat Claude’s Constitutional Classifiers in high-risk domains — a category distinct from ordinary software bugs, reflecting Anthropic’s stated view that model-behaviour failures are a security surface in their own right. Anthropic said the private phase of the programme had already produced findings that strengthened its products before the public launch.
The move follows a broader pattern among frontier labs of formalising external red-teaming as products and their misuse risks scale, rather than relying solely on internal testing.