Introducing a Comprehensive Dataset of Structured Biomedical Abstracts

This article presents a new dataset aimed at enhancing biomedical literature processing through structured abstracts.

3 min readBiomedical Research

Structured abstracts play a crucial role in the field of biomedical literature by improving the efficiency of information retrieval, text mining, and knowledge synthesis. Despite their importance, a significant number of abstracts in the PubMed database remain unstructured, creating challenges for various text-processing applications. To address this issue, we present Structured PubMed, a robust dataset containing section-labeled biomedical abstracts derived from the entire PubMed repository, which includes over 23.2 million research articles. The dataset is categorized into two main groups: one featuring 5.9 million abstracts that have been organized by authors from official XML files, and another comprising 17.2 million abstracts that were initially unstructured but have been systematically labeled using a Large Language Model extraction process. Each entry adheres to a standardized five-section format and is linked to its corresponding PubMed identifier, publication type, and date. This extensive dataset is designed to facilitate the training of sentence-classification models, evaluate text-segmentation systems, and enable large-scale, section-specific information extraction across the entire PubMed database.

Biomedical Research