ChatGPT gains voice and vision
A new text-to-speech model, built with professional voice actors, let users talk to ChatGPT and show it photos, ahead of the fully real-time GPT-4o release.
- Models & capabilities
- Notable
OpenAI began rolling out voice and image capabilities in ChatGPT, letting users speak to the assistant and receive a spoken reply, or share a photograph and ask questions about it. The features moved ChatGPT away from a purely text-based interface toward the fully multimodal, real-time product OpenAI would complete with GPT-4o eight months later.
The voice feature used a new text-to-speech model capable of generating natural-sounding speech from text and a short sample of a target voice, paired with Whisper, OpenAI’s existing open-source speech-recognition system, to transcribe what users said. OpenAI worked with professional voice actors to record five selectable voices rather than cloning any existing celebrity or public figure’s voice, a choice made ahead of the controversy a synthetic voice would later cause when GPT-4o’s launch drew comparisons to Scarlett Johansson.
Image understanding drew on GPT-4’s and GPT-3.5’s existing multimodal capabilities, letting the model reason over photographs, screenshots, and documents combining text and images. OpenAI’s examples — photographing a landmark while travelling and asking about it, or photographing a fridge’s contents to get recipe suggestions — emphasised everyday, low-stakes use over any specific professional application.
Access rolled out first to ChatGPT Plus and Enterprise subscribers, with voice arriving on iOS and Android and images available across platforms within two weeks of the announcement. The update tightened the field’s competitive frame: multimodality had until then been demonstrated mainly through research papers and GPT-4’s original launch materials, and this made speaking to, and showing things to, a chatbot into an ordinary consumer feature rather than a capability described in a technical report.