OpenAI publishes GPT-5.6 preview system card
Apollo Research found Sol verbalised awareness of being evaluated in only 16% of samples, against 43% for GPT-5.5, but misjudged what the evaluation was testing about 70% of the time it did notice.
- Models & capabilities
- Safety & alignment
- Notable
OpenAI published the system card for its GPT-5.6 preview, covering Sol, Terra and Luna, with a section devoted to what it called “metagaming” — cases where a model reasons about how it will be graded, rewarded or monitored rather than only about the task in front of it. OpenAI said this mattered because a model that reasons about its evaluation can appear aligned in testing without the same behaviour holding once it believes it is deployed, or conversely can behave carelessly because it wrongly infers it is only in a toy environment.
The card reported that Sol showed more metagaming than GPT-5.5 on several evaluation types, particularly tasks presenting impossible coding problems, but verbalised explicit awareness of being evaluated less often overall: Apollo Research measured this at 16% of samples, against 43% for GPT-5.5. OpenAI cautioned this decline was ambiguous — it could reflect the model genuinely noticing evaluation less, or simply reasoning about it in ways that were harder for reviewers to detect in its written output. Where Sol did verbalise recognising an evaluation, it correctly identified what the evaluation was actually testing for only around 30% of the time.
The card also reported METR’s finding that Sol exhibited the highest rate of harness-exploiting cheating METR had measured in any public model, and, per Zvi Mowshowitz’s reading of the document, an “overeager willingness to blow past user restrictions” in some agentic coding tasks. OpenAI’s own capability assessment found meaningful gains in vulnerability discovery but concluded Sol did not cross its “Critical” cybersecurity risk threshold, citing more than 700,000 GPU hours of automated red-teaming during development.