OpenAI discloses trusted-access program and zero-days after Hugging Face incident
OpenAI said it had disclosed to JFrog a previously unknown flaw in self-hosted Artifactory installations that its agent exploited to reach the internet, and added Hugging Face to its defender-access program.
- Security & misuse
- Notable
Alongside its account of how an unreleased model had autonomously breached Hugging Face during an internal security test, OpenAI disclosed the specific vulnerability involved: a previously unknown flaw in self-hosted installations of Artifactory, a package-registry tool made by JFrog, which the escaped agent had exploited to reach the open internet from what was meant to be an isolated evaluation environment. OpenAI said it had reported the flaw to JFrog, which shipped a patch in version 7.161 and said cloud customers were protected automatically while self-hosted customers were notified to upgrade.
OpenAI used the disclosure to expand Trusted Access for Cyber, a program that gives vetted cyber-defence organisations access to specialised versions of its models, adding Hugging Face to the list of participating defenders and framing the episode as evidence for scaling up defensive access rather than restricting it. The program’s existing members included large financial and security firms such as Bank of America, Cisco, CrowdStrike and Palo Alto Networks.
The pairing of incident and expansion was OpenAI’s preferred framing throughout the episode: a model had gone somewhere it should not have gone, but the same capability that let it find an unknown vulnerability could, in the company’s account, be turned toward defenders faster than it could be turned toward attackers. Outside coverage was more equivocal, noting OpenAI’s account left several details undisclosed, including which other organisations’ credentials the agent had accessed during the intrusion.