Fetching the paper…
Reading the bibliography…
Accurately extracting structured content from PDFs is a critical first step for NLP over scientific papers.
RoBERTAa: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Distilbert, a distilled version of BERT: Smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Random forests
Leo Breiman. 2001 · 2001
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
John D. Lafferty, Andrew McCallum, and Fernando C. N. Pereira. 2001 · 2001
Earlier work this paper cites.
CORD-19: the covid-19 open research dataset
Lucy Lu Wang, Kyle Lo, Yoganand Chandrasekhar, Russell Reas, Jiangjiang Yang, Darrin Eide, Kathryn Funk, Rodney Kinney, Ziyang Liu, William Merrill, Paul Mooney, Dewey A. Murdick, Devvret Rishi, Jerry Sheehan, Zhihong Shen, Brandon Stilson, Alex D. Wade, Kuansan Wang, Chris Wilhelm, Boya Xie, Douglas Raymond, Daniel S. Weld, Oren Etzioni, and Sebastian Kohlmeier. 2020 · 2004
Earlier work this paper cites.
GROTOAP2 — the methodology of creating a large ground truth dataset of scientific articles
Dominika Tkaczyk, Pawel Szostek, and Lukasz Bolikowski. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
CERMINE: Automatic extraction of structured metadata from scientific literature
Dominika Tkaczyk, Paweł Szostek, Mateusz Fedoryszak, Piotr Jan Dendek, and Łukasz Bolikowski. 2015 · 2015
Earlier work this paper cites.
Mask R-CNN
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Construction of the literature graph in semantic scholar
Waleed Ammar, Dirk Groeneveld, Chandra Bhagavatula, Iz Beltagy, Miles Crawford, Doug Downey, Jason Dunkelberger, Ahmed Elgohary, Sergey Feldman, Vu Ha, Rodney Kinney, Sebastian Kohlmeier, Kyle Lo, Tyler Murray, Hsu-Han Ooi, Matthew Peters, Joanna Power, Sam Skjonsberg, Lucy Wang, Chris Wilhelm, Zheng Yuan, Madeleine van Zuylen, and Oren Etzioni. 2018 · 2018
Earlier work this paper cites.
Mixed precision training
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory F. Diamos, Erich Elsen, David García, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu. 2018 · 2018
Cited alongside, same era.
Extracting scientific figures with distantly supervised neural networks
Noah Siegel, Nicholas Lourie, Russell Power, and Waleed Ammar. 2018 · 2018
Cited alongside, same era.
Corpus conversion service: A machine learning platform to ingest documents at scale
Peter WJ Staar, Michele Dolfi, Christoph Auer, and Costas Bekas. 2018 · 2018
Cited alongside, same era.
SciBERT: A pretrained language model for scientific text
Iz Beltagy, Kyle Lo, and Arman Cohan. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Green AI
Roy Schwartz, Jesse Dodge, Noah A Smith, and Oren Etzioni. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Later among the works it cites.
Layoutlm: Pre-training of text and layout for document image understanding
Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. 2020 · 2020
Later among the works it cites.
Beyond 512 tokens: Siamese multi-depth transformer-based hierarchical encoder for long-form document matching
Liu Yang, Mingyang Zhang, Cheng Li, Michael Bendersky, and Marc Najork. 2020 · 2020
Later among the works it cites.
ICDAR2021 competition on mathematical formula detection
2021 · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Cited alongside, same era.
PyTorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
Cited alongside, same era.
HIBERT: Document level pre-training of hierarchical bidirectional transformers for document summarization
Xingxing Zhang, Furu Wei, and Ming Zhou. 2019 · 2019
Cited alongside, same era.
PubLayNet: Largest dataset ever for document layout analysis
Xu Zhong, Jianbin Tang, and Antonio Jimeno Yepes. 2019 · 2019
Cited alongside, same era.
SLM: Learning a discourse language representation with sentence unshuffling
Haejun Lee, Drew A. Hudson, Kangwook Lee, and Christopher D. Manning. 2020 · 2020
Cited alongside, same era.
DocBank: A benchmark dataset for document layout analysis
Minghao Li, Yiheng Xu, Lei Cui, Shaohan Huang, Furu Wei, Zhoujun Li, and Ming Zhou. 2020 · 2020
Cited alongside, same era.
S2ORC: The semantic scholar open research corpus
Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel Weld. 2020 · 2020
Cited alongside, same era.
Segatron: Segment-aware transformer for language modeling and understanding
He Bai, Peng Shi, Jimmy Lin, Yuqing Xie, Luchen Tan, Kun Xiong, Wen Gao, and Ming Li. 2021 · 2021
Closest in time.
SelfDoc: Self-supervised document representation learning
Peizhao Li, Jiuxiang Gu, Jason Kuen, Vlad I Morariu, Handong Zhao, Rajiv Jain, Varun Manjunatha, and Hongfu Liu. 2021 · 2021
Closest in time.
Robust PDF document conversion using recurrent neural networks
Nikolaos Livathinos, Cesar Berrospi, Maksym Lysak, Viktor Kuropiatnyk, Ahmed Nassar, Andre Carvalho, Michele Dolfi, Christoph Auer, Kasper Dinkla, and Peter W. J. Staar. 2021 · 2021
Closest in time.
PAWLS: PDF annotation with labels and structure
Mark Neumann, Zejiang Shen, and Sam Skjonsberg. 2021 · 2021
Closest in time.
LayoutParser: A unified toolkit for deep learning based document image analysis
Zejiang Shen, Ruochen Zhang, Melissa Dell, Benjamin Charles Germain Lee, Jacob Carlson, and Weining Li. 2021 · 2021
Closest in time.
Lucy Lu Wang, Isabel Cachola, Jonathan Bragg, Evie Yu-Yen Cheng, Chelsea Haupt, Matt Latzke, Bailey Kuehl, Madeleine van Zuylen, Linda Wagner, and Daniel S. Weld. 2021 · 2021
Closest in time.
LayoutLMv2: Multi-modal pre-training for visually-rich document understanding
Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei A. F. Florêncio, Cha Zhang, Wanxiang Che, Min Zhang, and Lidong Zhou. 2021 · 2021
Closest in time.