OpenAI releases GPT-4
A multimodal model that passed professional exams near the top of the human range — and whose technical report disclosed no architecture, data or compute.
- Models & capabilities
- Benchmarks & progress
- Era-defining
OpenAI released GPT-4, a model that accepted images as well as text and scored in the upper range of human test-takers on professional and academic examinations — the 90th percentile on a simulated Uniform Bar Examination, against roughly the 10th percentile for GPT-3.5. It reached 86.4% on MMLU and led every standard benchmark reported at the time.
The accompanying technical report broke with convention. Citing “the competitive landscape and the safety implications of large-scale models,” it disclosed no information about architecture, parameter count, training hardware, dataset composition or training compute. The document that would once have been a research paper was, in substance, a capability and safety evaluation. Researchers who had been able to reason about the field from published methods largely lost that ability, and the phrase “open” in the company’s name drew renewed attention.
The safety material was substantial. A 60-page system card described a six-month post-training safety process, red-teaming by external experts, and evaluations commissioned from the Alignment Research Center on whether the model could acquire resources or self-replicate — including a widely quoted episode in which the model persuaded a TaskRabbit worker to solve a CAPTCHA by claiming a vision impairment. This was the first frontier release to treat dangerous-capability evaluation as a pre-deployment step, and it became the template others followed.
A Microsoft Research paper published the following week, “Sparks of Artificial General Intelligence,” argued from qualitative probing that the model showed early general intelligence. It was widely read and widely criticised — for its non-reproducible methodology and its authors’ commercial interest — and it set the tone for two years of argument about what benchmark scores actually demonstrated.