Fetching the paper…
Reading the bibliography…
Large-scale data sets on scholarly publications are the basis for a variety of bibliometric analyses and natural language processing (NLP) applications.
An index to quantify an individual’s scientific research output
Jorge E Hirsch. 2005 · 2005
Earlier work this paper cites.
CiteSpace II: Detecting and visualizing emerging trends and transient patterns in scientific literature
Chaomei Chen. 2006 · 2006
Earlier work this paper cites.
GROBID: Combining Automatic Bibliographic Data Recognition and Term Extraction for Scholarship Publications. In Research and Advanced Technology for Digital Libraries . 473–474
Patrice Lopez. 2009 · 2009
Earlier work this paper cites.
CITREC : An Evaluation Framework for Citation-Based Similarity Measures based on TREC Genomics and PubMed Central. In iConference 2015 Proceedings
Bela Gipp, Norman Meuschke, and Mario Lipinski. 2015 · 2015
Earlier work this paper cites.
An Overview of Microsoft Academic Service (MAS) and Applications. In Proceedings of the 24th International Conference on World Wide Web (Florence, Italy) (WWW’15) . 243–246
Arnab Sinha, Zhihong Shen, Yang Song, Hao Ma, Darrin Eide, Bo-June Paul Hsu, and Kuansan Wang. 2015 · 2015
Earlier work this paper cites.
Developing Infrastructure to Support Closer Collaboration of Aggregators with Open Repositories
Nancy Pontika, Petr Knoth, Matteo Cancellieri, and Samuel Pearce. 2016 · 2016
Earlier work this paper cites.
The FAIR Guiding Principles for scientific data management and stewardship
Mark D. Wilkinson, Michel Dumontier, IJsbrand Jan Aalbersberg, Gabrielle Appleton, Myles Axton, Arie Baak, Niklas Blomberg, Jan-Willem Boiten, Luiz Bonino da Silva Santos, Philip E. Bourne, Jildau Bouwman, Anthony J. Brookes, Tim Clark, Mercè Crosas, Ingrid Dillo, Olivier Dumon, Scott Edmunds, Chris T. Evelo, Richard Finkers, Alejandra Gonzalez-Beltran, Alasdair J. G. Gray, Paul Groth, Carole Goble, Jeffrey S. Grethe, Jaap Heringa, Peter A. C. ’t Hoen, Rob Hooft, Tobias Kuhn, Ruben Kok, Joost Kok, Scott J. Lusher, Maryann E. Martone, Albert Mons, Abel L. Packer, Bengt Persson, Philippe Rocca-Serra, Marco Roos, Rene van Schaik, Susanna-Assunta Sansone, Erik Schultes, Thierry Sengstag, Ted Slater, George Strawn, Morris A. Swertz, Mark Thompson, Johan van der Lei, Erik van Mulligen, Jan Velterop, Andra Waagmeester, Peter Wittenburg, Katherine Wolstencroft, Jun Zhao, and Barend Mons. 2016 · 2016
Earlier work this paper cites.
A Benchmark and Evaluation for Text Extraction from PDF. In 2017 ACM/IEEE Joint Conference on Digital Libraries (JCDL) . 1–10
H. Bast and C. Korzen. 2017 · 2017
Earlier work this paper cites.
Content-Based Citation Recommendation. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 . Association for Computational Linguistics, 238–251
Chandra Bhagavatula, Sergey Feldman, Russell Power, and Waleed Ammar. 2018 · 2018
Earlier work this paper cites.
Multi-Task Identification of Entities, Relations, and Coreferencefor Scientific Knowledge Graph Construction. In Proc. Conf. Empirical Methods Natural Language Process. (EMNLP)
Yi Luan, Luheng He, Mari Ostendorf, and Hannaneh Hajishirzi. 2018 · 2018
Cited alongside, same era.
Neural ParsCit: A Deep Learning Based Reference String Parser
Animesh Prasad, Manpreet Kaur, and Min-Yen Kan. 2018 · 2018
Cited alongside, same era.
Citation recommendation: approaches and datasets
Michael Färber and Adam Jatowt. 2020 · 2020
Cited alongside, same era.
arXMLiv:2020 dataset, an HTML5 conversion of arXiv.org
Deyan Ginev. 2020 · 2020
Cited alongside, same era.
S2ORC: The Semantic Scholar Open Research Corpus. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, 4969–4983
Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel Weld. 2020 · 2020
Cited alongside, same era.
SemEval-2021 Task 8: MeasEval – Extracting Counts and Measurements and their Related Contexts. In Proceedings of the 15th International Workshop on Semantic Evaluation (SemEval-2021) . 306–316
Corey Harper, Jessica Cox, Curt Kohler, Antony Scerri, Ron Daniel Jr., and Paul Groth. 2021 · 2021
Later among the works it cites.
Cross-Lingual Citations in English Papers: A Large-Scale Analysis of Prevalence, Usage, and Impact
Tarek Saier, Michael Färber, and Tornike Tsereteli. 2021 · 2021
Later among the works it cites.
CitationIE: Leveraging the Citation Graph for Scientific Information Extraction. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . 719–731
Vijay Viswanathan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Later among the works it cites.
SemEval 2022 Task 12: Symlink - Linking Mathematical Symbols to their Descriptions. In Proceedings of the 16th International Workshop on Semantic Evaluation (SemEval-2022) . 1671–1678
Viet Lai, Amir Pouran Ben Veyseh, Franck Dernoncourt, and Thien Nguyen. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
unarXive: a large scholarly data set with publications’ full-text, annotated in-text citations, and links to metadata
Tarek Saier and Michael Färber. 2020 · 2020
Cited alongside, same era.
Comparison of metadata with relevance for bibliometrics between Microsoft Academic Graph and OpenAlex until 2020
Thomas Scheidsteger and Robin Haunschild. 2022 · 2020
Cited alongside, same era.
Fact or Fiction: Verifying Scientific Claims. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) . 7534–7550
David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020 · 2020
Cited alongside, same era.
SciXGen: A Scientific Paper Dataset for Context-Aware Text Generation. In Findings of the Association for Computational Linguistics: EMNLP 2021 . 1483–1492
Hong Chen, Hiroya Takamura, and Hideki Nakayama. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
CiteSum: Citation Text-guided Scientific Extreme Summarization and Domain Adaptation with Limited Supervision. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing . 10922–10935
Yuning Mao, Ming Zhong, and Jiawei Han. 2022 · 2022
Later among the works it cites.
Multi-objective Representation Learning for Scientific Document Retrieval. In Proceedings of the Third Workshop on Scholarly Document Processing . Association for Computational Linguistics, Gyeongju, Republic of Korea, 80–88
Mathias Parisot and Jakub Zavrel. 2022 · 2022
Later among the works it cites.
OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts
Jason Priem, Heather Piwowar, and Richard Orr. 2022 · 2022
Later among the works it cites.
How Have Astronomers Cited Other Fields in the Last Decade?
Michele Delli Veneri, Rafael S. de Souza, Alberto Krone-Martins, E. E. O. Ishida, M. L. L. Dantas, Noble Kennamer, and for the COIN collaboration. 2022 · 2022
Later among the works it cites.
PMC Open Access Subset
Bethesda (MD): National Library of Medicine. [n. d.] · 2023
Closest in time.