The CORD-19 research dataset is released
Assembled with the White House OSTP, National Library of Medicine and Chan Zuckerberg Initiative, it opened with over 29,000 papers, most with full text.
- Ideas & essays
- Open weights & ecosystem
- Minor
The Allen Institute for AI, working with the White House Office of Science and Technology Policy, the National Library of Medicine, the Chan Zuckerberg Initiative, Microsoft Research, Kaggle and Georgetown’s Center for Security and Emerging Technology, released the first version of CORD-19: a machine-readable, full-text corpus of coronavirus and COVID-19 research literature. The initial release held over 29,000 articles, more than 13,000 of them with full text extracted rather than just an abstract.
The release accompanied a public “call to action,” in which the White House asked the AI and NLP research community to build text-mining and information-retrieval tools that could help researchers and clinicians surface answers buried in a literature growing faster than any one person could read. The dataset was formatted for machines rather than people — structured metadata plus parsed full text — specifically so search, question-answering and summarisation systems could be pointed at it directly.
CORD-19 continued to expand through 2020 as more publishers contributed content, and a companion paper describing the dataset’s construction and design choices was later posted to arXiv. It became a widely used benchmark and testbed for applied NLP work that spring, cited in hundreds of subsequent papers and downloaded well beyond the biomedical community it was built for, and stands as one of the clearer examples of the pandemic accelerating a specific, practical application of language technology.