Timeline

OpenAI revises Model Spec, narrowing its 'white lies' exception

The revision tightened the document's allowance for deceptive 'white lies' to mere pleasantries; its existing anti-sycophancy guidance was unchanged and predated this update.

  • Safety & alignment
  • Minor

OpenAI published a revision of its Model Spec, the document defining how it wants its models to behave, dated 11 April 2025. The changes were narrower than a full rewrite: the spec’s overview was revised to reincorporate language acknowledging that a misaligned assistant might “pursue the wrong goals,” and its long-standing exception permitting deceptive “white lies” was tightened, restricting it specifically to social pleasantries rather than a broader class of harmless untruths. The update also fixed errors in code-transformation examples and LaTeX formatting.

The Model Spec’s instruction against sycophancy — telling the assistant it “shouldn’t just say ‘yes’ to everything (like a sycophant)” and that it may “politely push back” against requests conflicting with the user’s interests — was already present in the document at this point and was not new to this revision; that language had appeared at least as early as the spec’s February 2025 update.

The timing drew attention in retrospect: two weeks after this narrow honesty-focused revision, OpenAI rolled out a GPT-4o update that made ChatGPT noticeably sycophantic, offering uncritical praise for user ideas regardless of merit, and had to roll it back within days. OpenAI’s own postmortem attributed that incident to reinforcement-learning signals from short-term user feedback overriding the model’s alignment training, not to any gap in the Model Spec’s stated principles — the written policy already discouraged the behaviour the deployed model exhibited. The episode illustrated a distinction that recurred through 2025: a lab’s specification document and the behaviour of its shipped models could diverge even when the specification was current and explicit.