Benchmarks · Language & multilingual

C-Eval

also: C-Eval Hard

How much a model knows and can reason about, tested in Chinese and calibrated to the Chinese education and professional-qualification system rather than translated from an English test.

Shanghai Jiao Tong University, Tsinghua University, University of Edinburgh & HKUSTReleased 15 May 2023Retired

C-Eval was built to answer a question that translated English benchmarks could not: how much does a model actually know when tested in Chinese, on material calibrated to Chinese schooling and professional exams rather than machine-translated from an American test bank? A consortium from Shanghai Jiao Tong University, Tsinghua University, the University of Edinburgh and HKUST assembled 13,948 multiple-choice questions across 52 subjects and four difficulty tiers, from middle school through professional qualification exams, along with a harder “C-Eval Hard” subset of subjects such as advanced mathematics that demand multi-step reasoning rather than recall.

At publication in mid-2023, the gap between English-centric and Chinese-native models was stark: GPT-4 was the only system to clear 60% average accuracy, reaching 66.4% in a zero-shot, answer-only setting, while other leading models of the day scored well behind it. That result, cited widely including in early releases from Chinese labs such as Alibaba’s Qwen family, made C-Eval a standard reference point for measuring how quickly Chinese-oriented open models were closing the gap with Western frontier systems on their own language.

The benchmark’s maintainers stopped updating its public leaderboard and released the full test set in mid-2025, a common endpoint for a benchmark once test-set exposure during training makes further comparisons unreliable. C-Eval’s questions and design also fed directly into Belebele-style and CMMLU-style successors that broadened Chinese-language evaluation beyond a single fixed question bank.

The set

13,948 multiple-choice questions across 52 subjects, split into four difficulty tiers — middle school, high school, college and professional — plus 'C-Eval Hard', a subset of the most reasoning-heavy subjects such as advanced mathematics and physics.

Example

'At 25°C, a pH=2 strong acid solution is mixed with a pH=13 strong base solution, giving a mixed solution of pH=11. What is the volume ratio of acid to base (ignoring volume change)?' Four options are given; the correct answer is 9:1.github.com

Where it stands

The maintainers stopped updating the public leaderboard in mid-2025 and released the full test set, so scores from strong Chinese-tuned models after that point should be treated cautiously given possible test-set exposure during training.

How the top score changed hands

  1. May 2023GPT-466.4%Zero-shot, answer-only setting; the only model in the paper's evaluation to pass 60% average accuracy.

In the timeline · 1 entry

More language & multilingual benchmarks