OpenAI launches Voice Engine research preview and its safety approach
OpenAI said it had held the text-to-speech model, first developed in 2022, back from wider release specifically because of election-year impersonation risk.
- Models & capabilities
- Safety & alignment
- Security & misuse
- Notable
OpenAI disclosed the existence of Voice Engine, a text-to-speech model able to generate a synthetic copy of a person’s voice from a single 15-second audio sample and a text prompt, while declining to release it beyond a small group of trusted partners. The company said the underlying model had been developed since late 2022 and already powered narrated features in ChatGPT’s voice mode and in Read Aloud tools, but that it was withholding the capability that had drawn most attention — cloning an arbitrary person’s voice from a short sample — from wider release.
OpenAI framed the decision explicitly around timing: it said the risks of voice-cloning technology were “especially top of mind” in a year with major elections across many countries, citing the potential for fraudulent impersonation and election interference. Partners given early access to Voice Engine were required to agree not to impersonate a person without that person’s explicit consent and to disclose that any voice a listener heard was AI-generated, and OpenAI said it embedded watermarking in Voice Engine’s audio output to help trace its origin.
The announcement was unusual in structure: rather than a product launch, it doubled as a policy statement, laying out mitigations OpenAI said it wanted the wider industry to adopt — including phasing out voice-based authentication for banking and other sensitive systems, since a cloned voice could defeat it — before OpenAI or any competitor released cloning tools broadly. It positioned OpenAI as flagging a capability gap between what its own research had already achieved and what it judged safe to ship, a stance it revisited seven weeks later when it paused a different voice feature, “Sky,” after Scarlett Johansson said it closely resembled her own voice despite her having declined to license it. Voice Engine itself remained in limited preview rather than general release through the rest of 2024.