Fetching the paper…
Reading the bibliography…
ClueWeb22, the newest iteration of the ClueWeb line of datasets, provides 10 billion web pages affiliated with rich information.
The PageRank Citation Ranking : Bringing Order to the Web. In Proceedings of WebConf 1999
Lawrence Page, Sergey Brin, Rajeev Motwani, and Terry Winograd. 1999 · 1999
Earlier work this paper cites.
ClueWeb09
Jamie Callan and Mark Hoy. 2009 · 2009
Earlier work this paper cites.
Overview of the TREC 2009 Web Track. In Proceedings of TREC 2009
Charles L. A. Clarke, Nick Craswell, and Ian Soboroff. 2009 · 2009
Earlier work this paper cites.
Building enriched document representations using aggregated anchor text. In Proceedings of SIGIR 2009
Donald Metzler, Jasmine Novak, Hang Cui, and Srihari Reddy. 2009 · 2009
Earlier work this paper cites.
ClueWeb12
Jamie Callan and David Pane. 2012 · 2012
Earlier work this paper cites.
Overview of the TREC 2012 Web Track. In Proceedings of TREC 2012
Charles L. A. Clarke, Nick Craswell, and Ellen M. Voorhees. 2012 · 2012
Earlier work this paper cites.
Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books. In Proceedings of IEEE 2015
Yukun Zhu, Ryan Kiros, Richard S. Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, et al · 2016
Earlier work this paper cites.
Reading Wikipedia to Answer Open-Domain Questions. In Proceedings of ACL 2017
Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017 · 2017
Earlier work this paper cites.
news-please: A Generic News Crawler and Extractor. In Proceedings of the ISI 2017
Felix Hamborg, Norman Meuschke, Corinna Breitinger, and Bela Gipp. 2017 · 2017
Earlier work this paper cites.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference. In Proceedings of NAACL-HLT 2018
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of NAACL-HLT 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
OpenWebText Corpus
Aaron Gokaslan and Vanya Cohen. 2019 · 2019
Cited alongside, same era.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Cited alongside, same era.
BlingFire
Microsoft. 2019 · 2019
Cited alongside, same era.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Open Domain Web Keyphrase Extraction Beyond Language Modeling. In Proceedings of EMNLP-IJCNLP 2019
Lee Xiong, Chuan Hu, Chenyan Xiong, Daniel Campos, and Arnold Overwijk. 2019 · 2019
Cited alongside, same era.
XLNet: Generalized Autoregressive Pretraining for Language Understanding. In Proceedings of NIPS 2019
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 2019
CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data. In Proceedings of LREC 2020
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave. 2020 · 2020
Later among the works it cites.
Selective Weak Supervision for Neural Information Retrieval. In Proceedings of WebConf 2020
Kaitao Zhang, Chenyan Xiong, Zhenghao Liu, and Zhiyuan Liu. 2020 · 2020
Later among the works it cites.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
William Fedus, Barret Zoph, and Noam Shazeer. 2021 · 2021
Later among the works it cites.
Open Domain Question Answering over Virtual Documents: A Unified Approach for Data and Text
Kaixin Ma, Hao Cheng, Xiaodong Liu, Eric Nyberg, and Jianfeng Gao. 2021 · 2021
Later among the works it cites.
COCO-LM: Correcting and Contrasting Text Sequences for Language Model Pretraining. In Proceedings of NIPS 2021
Yu Meng, Chenyan Xiong, Payal Bajaj, Saurabh Tiwary, Paul Bennett, Jiawei Han, and Xia Song. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
ETC: Encoding Long and Structured Inputs in Transformers. In Proceedings of EMNLP 2020
Joshua Ainslie, Santiago Ontañón, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, and Li Yang. 2020 · 2020
Cited alongside, same era.
Unsupervised Cross-lingual Representation Learning at Scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Cited alongside, same era.
The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, and Noa Nabeshima others. 2020 · 2020
Cited alongside, same era.
Retrieval Augmented Language Model Pre-Training. In Proceedings of ICML 2020
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020 · 2020
Cited alongside, same era.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Proceedings of NIPS 2020
Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Cited alongside, same era.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
From Web Graphs to Prioritizing Web Crawls
Sebastian Nage. 2021 · 2021
Later among the works it cites.
The Web Is Your Oyster - Knowledge-Intensive NLP against a Very Large Web Corpus
Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Dmytro Okhonko, Samuel Broscheit, Gautier Izacard, Patrick Lewis, Barlas Oguz, Edouard Grave, Wen-tau Yih, et al · 2021
Later among the works it cites.
Scaling Language Models: Methods, Analysis & Insights from Training Gopher
Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, H. Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al · 2021
Later among the works it cites.
PaLM: Scaling Language Modeling with Pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Closest in time.
GLaM: Efficient Scaling of Language Models with Mixture-of-Experts
Nan Du, Yanping Huang, Andrew M. Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, et al · 2022
Closest in time.
Training Compute-Optimal Large Language Models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Closest in time.