Organisation
UC Berkeley
A leading academic source of AI methods, benchmarks and researchers, home to the Berkeley AI Research (BAIR) lab.
The University of California, Berkeley is one of the leading academic homes of AI research, and its work recurs throughout the field's recent history. Berkeley researchers produced the MMLU benchmark that became the default test of a model's general knowledge, the Chatbot Arena that popularised ranking models by head-to-head human preference, and the vLLM inference engine widely used to serve large models efficiently. Much of this runs through the Berkeley AI Research lab, which brings together work across machine learning, vision, language and robotics. As a university rather than a company or government body, its role is to generate the methods, benchmarks and people that the rest of the field builds on, and it remains a major source of both.
- Category
- Academic labs & institutes
- Founded
- 1868
- HQ
- Berkeley, US
Appears alongside
Featured in threads
Tracks
- Benchmarks & progress 4
- Ideas & essays 2
- Safety & alignment 1
- Open weights & ecosystem 1
- Compute & infrastructure 1
UC Berkeley releases Agents' Last Exam, a benchmark of professional work
Built with 250+ industry experts across 55 sub-industries, it runs agents in the real software a specialist would use and grades against hidden answers; current systems clear under 1% of the hardest tier.
Benchmarks & progress
'Emergent Misalignment' shows narrow fine-tuning can broadly misalign a model
Fine-tuned only on insecure code with no disclosure of the flaws, GPT-4o and other models went on to endorse enslaving humanity and give malicious advice on unrelated prompts.
Ideas & essays · Safety & alignment
Scaling test-time compute paper argues extra inference compute can beat bigger models
On some problems, extra inference-time computation matched the gains from a pretrained model roughly 14 times larger, the authors reported.
Ideas & essays · Benchmarks & progress
vLLM releases PagedAttention inference engine
By managing attention memory in fixed-size pages rather than contiguous blocks, the UC Berkeley project cut memory waste and lifted serving throughput manyfold over existing stacks.
Open weights & ecosystem · Compute & infrastructure
LMSYS launches Chatbot Arena
The Berkeley-linked project ranked chatbots by anonymous, randomised head-to-head votes rather than fixed test sets, and later became LMArena.
Benchmarks & progress
Hendrycks et al. publish the MMLU benchmark
15,908-question, 57-subject multiple-choice benchmark spanning elementary to professional level; GPT-3 improved on random chance by roughly 20 points on average.
Benchmarks & progress