Fetching the paper…
Reading the bibliography…
Scientific literature serves as a high-quality corpus, supporting a lot of Natural Language Processing (NLP) research.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019 · 1910
Earlier work this paper cites.
Lcsts: A large scale chinese short text summarization dataset
Baotian Hu, Qingcai Chen, and Fangze Zhu. 2015 · 1972
Earlier work this paper cites.
Automatic evaluation of summaries using n-gram co-occurrence statistics
Chin-Yew Lin and Eduard Hovy. 2003 · 2003
Earlier work this paper cites.
Cluecorpus2020: A large-scale chinese corpus for pre-training language model
Liang Xu, Xuanwei Zhang, and Qianqian Dong. 2020 · 2003
Earlier work this paper cites.
Retrieval evaluation with incomplete information
Chris Buckley and Ellen M Voorhees. 2004 · 2004
Earlier work this paper cites.
Arnetminer: extraction and mining of academic social networks
Jie Tang, Jing Zhang, Limin Yao, Juanzi Li, Li Zhang, and Zhong Su. 2008 · 2008
Earlier work this paper cites.
Automatic keyphrase extraction from scientific articles
Su Nam Kim, Olena Medelyan, Min-Yen Kan, and Timothy Baldwin. 2013 · 2013
Earlier work this paper cites.
The acl anthology network corpus
Dragomir R Radev, Pradeep Muthukrishnan, Vahed Qazvinian, and Amjad Abu-Jbara. 2013 · 2013
Earlier work this paper cites.
An overview of microsoft academic service (mas) and applications
Arnab Sinha, Zhihong Shen, Yang Song, Hao Ma, Darrin Eide, Bo-June Hsu, and Kuansan Wang. 2015 · 2015
Earlier work this paper cites.
Paper recommender systems: a literature survey
Joeran Beel, Bela Gipp, Stefan Langer, and Corinna Breitinger. 2016 · 2016
Earlier work this paper cites.
Biocreative v cdr task corpus: a resource for chemical disease relation extraction
Jiao Li, Yueping Sun, Robin J Johnson, Daniela Sciaky, Chih-Hsuan Wei, Robert Leaman, Allan Peter Davis, Carolyn J Mattingly, Thomas C Wiegers, and Zhiyong Lu. 2016 · 2016
Cited alongside, same era.
Deep keyphrase generation
Rui Meng, Sanqiang Zhao, Shuguang Han, Daqing He, Peter Brusilovsky, and Yu Chi. 2017 · 2017
Cited alongside, same era.
Scientific document summarization via citation contextualization and scientific discourse
Arman Cohan and Nazli Goharian. 2018 · 2018
Cited alongside, same era.
A high-quality gold standard for citation-based tasks
Michael Färber, Alexander Thiemann, and Adam Jatowt. 2018 · 2018
Cited alongside, same era.
Measuring the evolution of a scientific field through citation frames
David Jurgens, Srijan Kumar, Raine Hoover, Dan McFarland, and Dan Jurafsky. 2018 · 2018
Cited alongside, same era.
Scibert: A pretrained language model for scientific text
Specter: Document-level representation learning using citation-informed transformers
Arman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey, and Daniel S Weld. 2020 · 2020
Later among the works it cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith. 2020 · 2020
Later among the works it cites.
Clts: A new chinese long text summarization dataset
Xiaojun Liu, Chuang Zhang, Xiaojun Chen, Yanan Cao, and Jinpeng Li. 2020 · 2020
Later among the works it cites.
S2orc: The semantic scholar open research corpus
Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel S Weld. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Iz Beltagy, Kyle Lo, and Arman Cohan. 2019 · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Bibliometric-enhanced arxiv: A data set for paper-based and citation-based tasks
Tarek Saier and Michael Färber. 2019 · 2019
Cited alongside, same era.
Oag: Toward linking large-scale heterogeneous entity graphs
Fanjin Zhang, Xiao Liu, Jie Tang, Yuxiao Dong, Peiran Yao, Jie Zhang, Xiaotao Gu, Yan Wang, Bin Shao, Rui Li, et al. 2019 · 2019
Cited alongside, same era.
Uer: An open-source toolkit for pre-training models
Zhe Zhao, Hui Chen, Jinbin Zhang, Wayne Xin Zhao, Tao Liu, Wei Lu, Xi Chen, Haotang Deng, Qi Ju, and Xiaoyong Du. 2019 · 2019
Cited alongside, same era.
unarxive: a large scholarly data set with publications’ full-text, annotated in-text citations, and links to metadata
Tarek Saier and Michael Färber. 2020 · 2020
Later among the works it cites.
Pegasus: Pre-training with extracted gap-sentences for abstractive summarization
Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu. 2020 · 2020
Later among the works it cites.
Flex: Unifying evaluation for few-shot nlp
Jonathan Bragg, Arman Cohan, Kyle Lo, and Iz Beltagy. 2021 · 2021
Later among the works it cites.
Crossfit: A few-shot learning challenge for cross-task generalization in nlp
Qinyuan Ye, Bill Yuchen Lin, and Xiang Ren. 2021 · 2021
Later among the works it cites.