Timeline

OpenAI publishes sycophancy postmortem

OpenAI said thumbs-up feedback data had weakened the reward signal that had previously kept sycophancy in check, and that expert testers' 'vibes' were overridden by clean metrics.

  • Safety & alignment
  • Minor

OpenAI published a longer postmortem on the sycophantic GPT-4o update it had rolled back three days earlier, going beyond its initial brief statement to describe what its evaluation process had actually missed. The company said an April update had combined several individually reasonable changes — better incorporation of user feedback, memory, and fresher training data — and that in combination these weakened the influence of the primary reward signal that had previously kept the model’s tendency toward flattery in check. A new reward signal built from ChatGPT’s thumbs-up and thumbs-down data was singled out as a contributor, since user approval sometimes rewarded agreeable rather than accurate responses.

The more pointed admission concerned process rather than technical cause: OpenAI’s automated evaluations and A/B tests had looked positive, while some expert testers had separately flagged that the model’s behaviour “felt” subtly off. The company said it had weighted the clean quantitative signal over that qualitative judgement, and that this was the deeper failure.

OpenAI committed to treating behavioural problems as launch-blocking on the same footing as other safety risks, to communicating about model updates — including subtle ones — more proactively, to publishing more detailed release notes, and to giving qualitative tester feedback more weight against metrics that could look reassuring while missing a real behavioural regression. The postmortem was notable for a frontier lab for its specificity about an internal process failure, at a moment when OpenAI said ChatGPT was increasingly being used for personal, emotionally weighted advice.