Fetching the paper…
Reading the bibliography…
Modern machine learning relies on datasets to develop and validate research ideas.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Measuring agreement for multinomial data
Mark Davies and Joseph L. Fleiss. 1982 · 1982
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz. 1993 · 1993
Earlier work this paper cites.
Okapi/Keenbow at TREC-8
Stephen E. Robertson and Steve Walker. 1999 · 1999
Earlier work this paper cites.
Optimizing search engines using clickthrough data
Thorsten Joachims. 2002 · 2002
Earlier work this paper cites.
Nltk: The natural language toolkit
Edward Loper and Steven Bird. 2002 · 2002
Earlier work this paper cites.
Evaluating web-based question answering systems
Dragomir R. Radev, Hong Qi, Harris Wu, and Weiguo Fan. 2002 · 2002
Earlier work this paper cites.
A statistical interpretation of term specificity and its application in retrieval
Karen Spärck Jones. 2004 · 2004
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
Citances: Citation sentences for semantic analysis of bioscience text
Preslav Nakov, Ariel S. Schwartz, and Marti A. Hearst. 2004 · 2004
Earlier work this paper cites.
Introduction to Information Retrieval
Christopher D. Manning, Prabhakar Raghavan, and Hinrich Schütze. 2005 · 2005
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, K. Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
The probabilistic relevance framework: BM25 and beyond
Stephen E. Robertson and Hugo Zaragoza. 2009 · 2009
Earlier work this paper cites.
The Pascal Visual Object Classes (VOC) challenge
Mark Everingham, Luc Van Gool, Christopher K. I. Williams, John M. Winn, and Andrew Zisserman. 2010 · 2010
Earlier work this paper cites.
Linguistic annotation
Martha Palmer and Nianwen Xue. 2010 · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012 · 2012
Earlier work this paper cites.
A dataset search engine for the research document corpus
Meiyu Lu, Srinivas Bangalore, Graham Cormode, Marios Hadjieleftheriou, and Divesh Srivastava. 2012 · 2012
Cited alongside, same era.
Vision meets robotics: The kitti dataset
Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun. 2013 · 2013
Cited alongside, same era.
Dataset retrieval
Sven R. Kunze and Sören Auer. 2013 · 2013
Cited alongside, same era.
Edinburgh SLT and MT system description for the iwslt 2014 evaluation
Alexandra Birch, Matthias Huck, Nadir Durrani, Nikolay Bogoychev, and Philipp Koehn. 2014 · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Cited alongside, same era.
Dataset recommendation for data linking: An intensional approach
Mohamed Ben Ellefi, Zohra Bellahsene, Stefan Dietze, and Konstantin Todorov. 2016 · 2016
SciREX: A challenge dataset for document-level information extraction
Sarthak Jain, Madeleine van Zuylen, Hannaneh Hajishirzi, and Iz Beltagy. 2020 · 2020
Later among the works it cites.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020 · 2020
Later among the works it cites.
S2ORC: The semantic scholar open research corpus
Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel Weld. 2020 · 2020
Later among the works it cites.
Recommending datasets for scientific problem descriptions
Michael Färber and Ann-Kathrin Leisinger. 2021 · 2021
Later among the works it cites.
Rethink training of BERT rerankers in multi-stage retrieval pipeline
Luyu Gao, Zhuyun Dai, and Jamie Callan. 2021 · 2021
Later among the works it cites.
Datasets: A community library for natural language processing
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ms marco: A human generated machine reading comprehension dataset
Daniel Fernando Campos, Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, Li Deng, and Bhaskar Mitra. 2016 · 2016
Cited alongside, same era.
Dats, the data tag suite to enable discoverability of datasets
Susanna-Assunta Sansone, Alejandra N. González-Beltrán, Philippe Rocca-Serra, George Alter, Jeffrey S. Grethe, Hua Xu, Ian M. Fore, Jared Lyle, Anupama E. Gururaj, Xiaoling Chen, Hyeon eui Kim, Nansu Zong, Yueling Li, Ruiling Liu, I. B. Ozyurt, and Lucila Ohno-Machado. 2017 · 2017
Cited alongside, same era.
An analysis of environment, microphone and data simulation mismatches in robust speech recognition
E. Vincent, S. Watanabe, A. Nugraha, J. Barker, and R. Marxer. 2017 · 2017
Cited alongside, same era.
Dataset recommendation via variational graph autoencoder
Basmah Altaf, Uchenna Akujuobi, Lu Yu, and Xiangliang Zhang. 2019 · 2019
Cited alongside, same era.
SciBERT: Pretrained language model for scientific text
Iz Beltagy, Kyle Lo, and Arman Cohan. 2019 · 2019
Cited alongside, same era.
Google Dataset Search: Building a search engine for datasets in an open web ecosystem
Dan Brickley, Matthew Burgess, and Natasha Noy. 2019 · 2019
Cited alongside, same era.
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario vSavsko, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clement Delangue, Th’eo Matussiere, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, Franccois Lagunas, Alexander M. Rush, and Thomas Wolf. 2021 · 2021
Later among the works it cites.
Pyserini: A Python Toolkit for Reproducible Information Retrieval Research with Sparse and Dense Representations
Jimmy Lin, Xueguang Ma, Sheng-Chieh Lin, Jheng-Hong Yang, Ronak Pradeep, and Rodrigo Nogueira. 2021 · 2021
Later among the works it cites.
Data and its (dis) contents: A survey of dataset development and use in machine learning research
Amandalynne Paullada, Inioluwa Deborah Raji, Emily M Bender, Emily Denton, and Alex Hanna. 2021 · 2021
Later among the works it cites.
Want to reduce labeling cost? gpt-3 can help
Shuohang Wang, Yang Liu, Yichong Xu, Chenguang Zhu, and Michael Zeng. 2021 · 2021
Later among the works it cites.
Counting AI research: Exploring ai research output in english- and chinese-language sources
Daniel Chou. 2022 · 2022
Later among the works it cites.
A systematic study of bias amplification
Melissa R.H. Hall, Laurens van der Maaten, Laura Gustafson, and Aaron B. Adcock. 2022 · 2022
Later among the works it cites.
Galactica: A large language model for science
Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic. 2022 · 2022
Later among the works it cites.
Evaluating token-level and passage-level dense retrieval models for math information retrieval
Wei Zhong, Jheng-Hong Yang, Yuqing Xie, and Jimmy J. Lin. 2022 · 2022
Later among the works it cites.
Semantic understanding of scenes through the ade20k dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fidler, Adela Barriuso, and Antonio Torralba. 2019 · 2022
Later among the works it cites.
Retrieving texts based on abstract descriptions
Shauli Ravfogel, Valentina Pyatkin, Amir D. N. Cohen, Avshalom Manevich, and Yoav Goldberg. 2023 · 2023
Closest in time.