2017

PubMed 200k RCT: a Dataset for Sequential Sentence Classification in Medical Abstracts

Dernoncourt, Franck, Lee, Ji Young

Understand

We present PubMed 200k RCT, a new dataset based on PubMed for sequential sentence classification.

  • The dataset consists of approximately 200,000 abstracts of randomized controlled trials, totaling 2.3 million sentences.
  • Each sentence of each abstract is labeled with their role in the abstract using one of the following classes: background, objective, method, result, or conclusion.
  • The purpose of releasing this dataset is twofold.

Reading the bibliography…