Timeline

OpenAI publishes the GPT-4V(ision) system card

A three-month alpha and over 50 outside experts probing high-risk areas informed mitigations against uses such as CAPTCHA-solving and identifying people from photos.

  • Safety & alignment
  • Models & capabilities
  • Minor

OpenAI published a system card documenting the safety evaluation behind GPT-4V(ision), a version of GPT-4 able to accept images alongside text, ahead of the capability’s wider rollout to ChatGPT Plus subscribers. The underlying model had finished training in 2022; OpenAI said it had spent the months since running a limited early-access period, deployed to over 1,000 alpha users, specifically to surface safety problems before a broader release.

More than 50 external red-teamers with expertise in areas including medicine, finance and cybersecurity tested the model for high-risk misuse. Two findings shaped the mitigations OpenAI shipped. First, the model could identify people from photographs with some reliability, so OpenAI built both a refusal evaluation, testing whether the model would decline person-identification requests, and an accuracy evaluation of how often it was right when it complied; the deployed system was tuned to refuse most such requests. Second, testers found the model could solve CAPTCHA challenges when directed to, a capability OpenAI judged had broader cybersecurity implications if made easily accessible, and trained the model to decline.

The card also disclosed unresolved limitations rather than only mitigations. OpenAI said the model remained unreliable for high-stakes scientific tasks such as identifying dangerous chemical compounds, and that it could hallucinate or misread images — for instance, merging two pieces of text that appeared close together in an image into one. Medical-imaging testing found the model prone to factual errors and inconsistent interpretations when given insufficient context, and OpenAI recommended against relying on it for diagnosis.

The document extended the practice the original GPT-4 system card had established six months earlier — publishing a capability and risk evaluation in place of technical detail about the model itself — into a second, imaging-specific system card, reinforcing pre-deployment safety testing as a step OpenAI treated as routine for each major new capability rather than a one-off exercise tied to GPT-4’s initial release.