Google Health reports AI matching radiologists on breast screening
A Nature paper claimed an AI system reduced false positives and false negatives in mammography against expert readers.
- Models & capabilities
- Benchmarks & progress
- Notable
Researchers from Google Health and DeepMind, working with Cancer Research UK’s Imperial Centre, Northwestern University and Royal Surrey County Hospital, published a retrospective study in Nature reporting that an AI system reduced both false positives and false negatives in mammogram reading compared with individual radiologists. On a UK dataset of 25,856 women screened between 2012 and 2015, the system cut false positives by 1.2% and false negatives by 2.7%; on a smaller US dataset of 3,097 women screened between 2001 and 2018, the reductions were larger — 5.7% and 9.4% respectively. The system worked from mammogram images alone, without the prior scans and clinical history that the human readers had access to.
The study was retrospective — the system was tested against archived cases with known outcomes, not deployed on new patients — and the authors were explicit that prospective clinical trials would be needed before any claim about real-world performance could be made. It arrived amid a run of similar “AI beats doctors” results across radiology, and this one drew unusually sustained scrutiny: in October 2020, Nature published a formal comment from a group of researchers led by Benjamin Haibe-Kains arguing that the paper’s description of its methods and code was too sparse to allow independent replication, undermining its value as a scientific contribution regardless of the headline numbers. McKinney and colleagues published a reply and an addendum expanding their supplementary methods, but did not release the model or training code.
The episode became a reference point in a broader argument about reproducibility in AI-for-medicine research — that a benchmark claim published in a leading journal is not, by itself, evidence a system works outside the conditions it was tested under, and that transparency about method matters as much as the result.