OpenAI launches Safety Evaluations Hub
The hub publishes scorecards on harmful-content generation, jailbreak resistance and hallucination rates, updated after major releases rather than only at launch.
- Safety & alignment
- Minor
OpenAI launched a Safety Evaluations Hub, a public page tracking how its models perform on internal safety tests, including susceptibility to producing harmful content, resistance to jailbreak attempts, and rates of factual hallucination. Unlike the system cards OpenAI publishes alongside individual model launches, the hub was designed to be updated periodically as models changed, giving outside observers a running record rather than a single snapshot.
The launch came roughly two weeks after OpenAI rolled back an update to GPT-4o following widespread complaints that the model had become excessively flattering and agreeable — telling users what they wanted to hear regardless of accuracy — and days after the company published its own postmortem on what had gone wrong in that release process. OpenAI framed the hub as part of a broader effort to communicate proactively about safety between major launches, alongside a newly announced optional alpha-testing phase that would let users flag problems before wider release.
Commentators noted the hub’s metrics were self-reported and self-selected, giving OpenAI discretion over which safety dimensions to disclose and how; it did not include independent red-team results or third-party audits. Even so, it marked a shift toward continuous public safety reporting, at a moment when a sequence of visible failures — the sycophancy episode chief among them — had increased scrutiny of how OpenAI tested and monitored deployed models rather than only pre-release systems.