Fetching the paper…
Reading the bibliography…
Large datasets have become commonplace in NLP research.
Learning and evaluating general linguistic intelligence
Dani Yogatama, Cyprien de Masson d’Autume, Jerome Connor, Tomas Kocisky, Mike Chrzanowski, Lingpeng Kong, Angeliki Lazaridou, Wang Ling, Lei Yu, Chris Dyer, et al. 2019 · 1901
Earlier work this paper cites.
Evaluating scalable bayesian deep learning methods for robust computer vision
Fredrik K Gustafsson, Martin Danelljan, and Thomas B Schön. 2019 · 1906
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar S. Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke S. Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R’emi Louf, Morgan Funtowicz, and Jamie Brew. 2019 · 1910
Earlier work this paper cites.
Deep ensembles: A loss landscape perspective
Stanislav Fort, Huiyi Hu, and Balaji Lakshminarayanan. 2019 · 1912
Earlier work this paper cites.
olmpics - on what language model pre-training captures
Alon Talmor, Yanai Elazar, Yoav Goldberg, and Jonathan Berant. 2019 · 1912
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
A sequential algorithm for training text classifiers
David D. Lewis and William A. Gale. 1994 · 1994
Earlier work this paper cites.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah A. Smith. 2020 · 2002
Earlier work this paper cites.
Distinguishing easy and hard instances
Yuval Krymolowski. 2002 · 2002
Earlier work this paper cites.
Pretrained transformers improve out-of-distribution robustness
Dan Hendrycks, Xiaoyuan Liu, Eric Wallace, Adam Dziedzic, Rishabh Krishnan, and Dawn Song. 2020 · 2004
Earlier work this paper cites.
Breaking SVM complexity with Cross-Training
Léon Bottou, Jason Weston, and Gökhan H Bakir. 2005 · 2005
Earlier work this paper cites.
Get another label? improving data quality and data mining using multiple, noisy labelers
Victor S Sheng, Foster Provost, and Panagiotis G Ipeirotis. 2008 · 2008
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. 2009 · 2009
Earlier work this paper cites.
Multi-class active learning for image classification
Ajay J Joshi, Fatih Porikli, and Nikolaos Papanikolopoulos. 2009 · 2009
Earlier work this paper cites.
Active learning literature survey
Burr Settles. 2009 · 2009
Earlier work this paper cites.
Self-paced learning for latent variable models
M Pawan Kumar, Benjamin Packer, and Daphne Koller. 2010 · 2010
Earlier work this paper cites.
Learning the easy things first: Self-paced visual category discovery
Yong Jae Lee and Kristen Grauman. 2011 · 2011
Earlier work this paper cites.
The Winograd schema challenge
Hector J Levesque, Ernest Davis, and Leora Morgenstern. 2011 · 2011
Earlier work this paper cites.
Part-of-speech tagging from 97% to 100%: Is it time for some linguistics?
Christopher D. Manning. 2011 · 2011
Earlier work this paper cites.
Unbiased look at dataset bias
Antonio Torralba and Alexei A Efros. 2011 · 2011
Earlier work this paper cites.
Facility location: concepts, models, algorithms and case studies. series: Contributions to management science
Gert W Wolf. 2011 · 2011
Earlier work this paper cites.
Learning whom to trust with MACE
Dirk Hovy, Taylor Berg-Kirkpatrick, Ashish Vaswani, and Eduard Hovy. 2013 · 2013
Earlier work this paper cites.
Using document summarization techniques for speech data subset selection
Kai Wei, Yuzong Liu, Katrin Kirchhoff, and Jeff Bilmes. 2013 · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Cited alongside, same era.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. 2015 · 2015
Cited alongside, same era.
Uncertainty in Deep Learning
Yarin Gal. 2016 · 2016
Cited alongside, same era.
Are all training examples created equal? an empirical study
Kailas Vodrahalli, Ke Li, and Jitendra Malik. 2018 · 2018
Later among the works it cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Later among the works it cites.
Chen Xing, Devansh Arpit, Christos Tsirigotis, and Yoshua Bengio. 2018 · 2018
Later among the works it cites.
Unsupervised label noise modeling and loss correction
Eric Arazo, Diego Ortego, Paul Albert, Noel O’Connor, and Kevin Mcguinness. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani. 2016 · 2016
Cited alongside, same era.
Embracing error to enable rapid crowdsourcing
Ranjay A. Krishna, Kenji Hata, Stephanie Chen, Joshua Kravitz, David A. Shamma, Fei-Fei Li, and Michael S. Bernstein. 2016 · 2016
Cited alongside, same era.
Online batch selection for faster training of neural networks
Ilya Loshchilov and Frank Hutter. 2016 · 2016
Cited alongside, same era.
SQuAD: 100, 000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
Training region-based object detectors with online hard example mining
Abhinav Shrivastava, Abhinav Gupta, and Ross Girshick. 2016 · 2016
Cited alongside, same era.
Active bias: Training more accurate neural networks by emphasizing high variance samples
Haw-Shiuan Chang, Erik Learned-Miller, and Andrew McCallum. 2017 · 2017
Cited alongside, same era.
Finding label noise examples in large scale datasets
R. Ekambaram, D. B. Goldgof, and L. O. Hall. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
Understanding and utilizing deep neural networks trained with noisy labels
Pengfei Chen, Benben Liao, Guangyong Chen, and Shengyu Zhang. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Show your work: Improved reporting of experimental results
Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A. Smith. 2019 · 2019
Later among the works it cites.
Data shapley: Equitable valuation of data for machine learning
Amirata Ghorbani and James Zou. 2019 · 2019
Later among the works it cites.
REPAIR: Removing representation bias by dataset resampling
Yi Li and Nuno Vasconcelos. 2019 · 2019
Later among the works it cites.
Data-efficient neural text compression with interactive learning
Avinesh P.V.S and Christian M. Meyer. 2019 · 2019
Later among the works it cites.
Learning with bad training data via iterative trimmed loss minimization
Yanyao Shen and Sujay Sanghavi. 2019 · 2019
Later among the works it cites.
Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift
Jasper Snoek, Yaniv Ovadia, Emily Fertig, Balaji Lakshminarayanan, Sebastian Nowozin, D Sculley, Joshua Dillon, Jie Ren, and Zachary Nado. 2019 · 2019
Later among the works it cites.
SuperGLUE: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Later among the works it cites.
Wei Hu, Zhiyuan Li, and Dingli Yu. 2020 · 2020
Closest in time.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2020
Closest in time.
Adversarial filters of dataset biases
Ronan LeBras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers, Matthew E. Peters, Ashish Sabharwal, and Yejin Choi. 2020 · 2020
Closest in time.
How can we accelerate progress towards human-like linguistic generalization?
Tal Linzen. 2020 · 2020
Closest in time.
DQI: Measuring Data Quality in NLP
Swaroop Mishra, Anjana Arunkumar, Bhavdeep Sachdeva, Chris Bryan, and Chitta Baral. 2020 · 2020
Closest in time.
Continual deep learning by functional regularisation of memorable past
Pingbo Pan, Siddharth Swaroop, Alexander Immer, Runa Eschenhagen, Richard E. Turner, and Mohammad Emtiyaz Khan. 2020 · 2020
Closest in time.
Identifying mislabeled data using the area under the margin ranking
Geoff Pleiss, Tianyi Zhang, Ethan R. Elenberg, and Kilian Q. Weinberger. 2020 · 2020
Closest in time.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020 · 2020
Closest in time.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2031
Closest in time.