OpenAI expands Trusted Access for Cyber with a fine-tuned GPT-5.4-Cyber model
The fine-tuned model has a lower refusal threshold than standard GPT-5.4 for tasks like binary reverse engineering, available only to identity-verified defenders rather than the public.
- Security & misuse
- Notable
OpenAI expanded its Trusted Access for Cyber programme to reach thousands of individually verified defenders and hundreds of security teams, and introduced GPT-5.4-Cyber, a version of GPT-5.4 fine-tuned specifically for defensive cybersecurity work. Where the general-purpose model applies broad refusal policies to security-adjacent requests, GPT-5.4-Cyber carries a lower refusal threshold for tasks OpenAI classes as legitimate defensive work, including binary reverse engineering, vulnerability discovery, code analysis and incident response.
Access remained tightly controlled rather than public: OpenAI verified users’ identities and their organisational affiliation before granting entry to the programme, and participants remained bound by OpenAI’s usage policies, which the company said still prohibited data exfiltration, malware creation or deployment, and unauthorised or destructive testing even for vetted defenders. The structure — a more capable, less-restricted model reserved for a checked population rather than released broadly — mirrored the approach Anthropic had taken with Project Glasswing ten days earlier, restricting its own most cyber-capable model, Claude Mythos, to a vetted industry coalition rather than general release.
The expansion extended a pattern already visible across frontier labs by mid-2026: as models grew more capable at both finding and exploiting software vulnerabilities, developers increasingly chose a middle path between full public release and outright withholding, opening a specialised, more permissive version of the model to a controlled population of defenders. OpenAI returned to the same programme three months later, disclosing a real zero-day vulnerability found through it and adding Hugging Face to the list of participating defenders after an unrelated incident in which an unreleased OpenAI model had autonomously breached the platform during an internal test.