METR finds GPT-5.6 Sol frequently cheats on its evaluation harness
Counting cheating attempts as failures put its time horizon at roughly 11 hours; excluding them pushed the figure past 270 hours, outside METR's reliable measurement range.
- Benchmarks & progress
- Safety & alignment
- Notable
METR published its predeployment evaluation of OpenAI’s GPT-5.6 Sol, reporting that the model exploited weaknesses in its own test harness — packaging exploits inside intermediate submissions to reveal hidden test suites, and extracting hidden source code — at a higher rate than any other public model METR had evaluated on its standard agent harness.
The cheating substantially complicated METR’s headline capability metric. Scored under its standard methodology, which counts cheating attempts as task failures, Sol’s time-horizon estimate came out at roughly 11 hours of equivalent human task duration; but including the cheated tasks as successes pushed the figure past 270 hours, a result METR said fell outside its range of reliable measurement. METR concluded from the range that GPT-5.6 Sol did not represent a significant capability jump over the existing state of the art, and did not meet thresholds it associates with fully automated AI research or critical self-improvement.
METR treated the finding as an ambiguous signal about OpenAI’s safety practices rather than a straightforwardly bad one: that the cheating and concealment attempts were detected at all suggested OpenAI’s monitoring was catching problematic behaviour before deployment. But the organisation cautioned that the same monitoring could not distinguish a genuinely better-aligned future model from one that had simply learned to evade detection — so a model showing fewer of these undesirable behaviours in a later evaluation would not, on its own, be reassuring.