Google Research launches an AI co-scientist to help generate hypotheses
A multi-agent Gemini 2.0 system with Generation, Reflection and Ranking agents proposed drug candidates later confirmed active in laboratory tests for two diseases.
- Models & capabilities
- Notable
Google Research announced an “AI co-scientist,” a multi-agent system built on Gemini 2.0 designed to work alongside researchers as a collaborator that proposes novel hypotheses and experimental plans, rather than as a single question-answering model.
The system split the work across specialised agents — Generation, Reflection, Ranking, Evolution, Proximity and Meta-review, coordinated by a Supervisor agent — modelled loosely on stages of the scientific method: generating candidate ideas, critiquing them, running them through a tournament-style ranking process, and refining the survivors, with web search and specialised models used to ground hypotheses in existing evidence. Google described this as a form of test-time compute scaling, in which the system spends additional inference-time computation on self-play debate and iterative refinement rather than relying only on what was learned during training.
Google reported validating the system with outside collaborators across three biomedical problems: identifying drug-repurposing candidates for acute myeloid leukaemia, proposing epigenetic targets for liver fibrosis, and explaining a bacterial mechanism of antimicrobial resistance. Domain experts judged the AI co-scientist’s hypotheses to have higher potential novelty and impact than those from baseline models, and several proposed drugs were subsequently tested in the laboratory; a liver-fibrosis follow-up study published two months later reported that two of three AI-recommended epigenetic-modifier drugs showed significant anti-fibrotic activity in tissue models.
Rather than a general release, Google opened access through a “Trusted Tester Program,” inviting research institutions to apply to evaluate the system on their own problems ahead of any wider rollout. The launch extended a push, running from AlphaFold onward, to move large models from generating text toward acting as active participants earlier in the research process; the system’s output was later put through formal peer review, culminating in a Nature paper in 2026.