Fetching the paper…
Reading the bibliography…
Models that top leaderboards often perform unsatisfactorily when deployed in real world applications; this has necessitated rigorous and expensive pre-deployment model testing.
Do imagenet classifiers generalize to imagenet?
Recht, B.; Roelofs, R.; Schmidt, L.; and Shankar, V. 2019 · 1902
Earlier work this paper cites.
Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment
Jin, D.; Jin, Z.; Tianyi Zhou, J.; and Szolovits, P. 2019 · 1907
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K.; Bras, R. L.; Bhagavatula, C.; and Choi, Y. 2019 · 1907
Earlier work this paper cites.
Schwartz, R.; Dodge, J.; Smith, N. A.; and Etzioni, O. 2019 · 1907
Earlier work this paper cites.
Perturbation sensitivity analysis to detect unintended model biases
Prabhakaran, V.; Hutchinson, B.; and Mitchell, M. 2019 · 1910
Earlier work this paper cites.
Distributional structure
Harris, Z. S. 1954 · 1954
Earlier work this paper cites.
Black-box testing: techniques for functional testing of software and systems
Beizer, B. 1995 · 1995
Earlier work this paper cites.
Convolutional networks for images, speech, and time series
LeCun, Y.; Bengio, Y.; et al. 1995 · 1995
Earlier work this paper cites.
Long short-term memory
Hochreiter, S.; and Schmidhuber, J. 1997 · 1997
Earlier work this paper cites.
Adversarial Filters of Dataset Biases
Bras, R. L.; Swayamdipta, S.; Bhagavatula, C.; Zellers, R.; Peters, M. E.; Sabharwal, A.; and Choi, Y. 2020 · 2002
Earlier work this paper cites.
Towards the Systematic Reporting of the Energy and Carbon Footprints of Machine Learning
Henderson, P.; Hu, J.; Romoff, J.; Brunskill, E.; Jurafsky, D.; and Pineau, J. 2020 · 2002
Earlier work this paper cites.
Evaluation challenges in large-scale document summarization
Radev, D.; Teufel, S.; Saggion, H.; Lam, W.; Blitzer, J.; Qi, H.; Celebi, A.; Liu, D.; and Drabek, E. F. 2003 · 2003
Earlier work this paper cites.
Pretrained Transformers Improve Out-of-Distribution Robustness
Hendrycks, D.; Liu, X.; Wallace, E.; Dziedzic, A.; Krishnan, R.; and Song, D. 2020 · 2004
Earlier work this paper cites.
DQI: Measuring Data Quality in NLP
Mishra, S.; Arunkumar, A.; Sachdeva, B. S.; Bryan, C.; and Baral, C. 2020b · 2005
Earlier work this paper cites.
Selective Question Answering under Domain Shift
Kamath, A.; Jia, R.; and Liang, P. 2020 · 2006
Cited alongside, same era.
Our Evaluation Metric Needs an Update to Encourage Generalization
Mishra, S.; Arunkumar, A.; Bryan, C.; and Baral, C. 2020a · 2007
Cited alongside, same era.
Question and Answer Test-Train Overlap in Open-Domain Question Answering Datasets
Lewis, P.; Stenetorp, P.; and Riedel, S. 2020 · 2008
Cited alongside, same era.
Human-robot interaction and future industrial robotics applications
Heyer, C. 2010 · 2010
Cited alongside, same era.
Learning word vectors for sentiment analysis
Maas, A. L.; Daly, R. E.; Pham, P. T.; Huang, D.; Ng, A. Y.; and Potts, C. 2011 · 2011
Cited alongside, same era.
Synthetic and natural noise both break neural machine translation
Belinkov, Y.; and Bisk, Y. 2017 · 2017
Later among the works it cites.
Adversarial Examples for Evaluating Reading Comprehension Systems
Jia, R.; and Liang, P. 2017 · 2017
Later among the works it cites.
Reimers, N.; and Gurevych, I. 2017 · 2017
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Later among the works it cites.
Adversarial example generation with syntactically controlled paraphrase networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Unbiased look at dataset bias
Torralba, A.; and Efros, A. A. 2011 · 2011
Cited alongside, same era.
Assessing the influence of climate model uncertainty on EU-wide climate change impact indicators
Lung, T.; Dosio, A.; Becker, W.; Lavalle, C.; and Bouwer, L. M. 2013 · 2013
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Mikolov, T.; Sutskever, I.; Chen, K.; Corrado, G. S.; and Dean, J. 2013 · 2013
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R.; Perelygin, A.; Wu, J.; Chuang, J.; Manning, C. D.; Ng, A. Y.; and Potts, C. 2013 · 2013
Cited alongside, same era.
Projection and uncertainty analysis of global precipitation-related extremes using CMIP5 models
Chen, H.; Sun, J.; and Chen, X. 2014 · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
Pennington, J.; Socher, R.; and Manning, C. D. 2014 · 2014
Cited alongside, same era.
The ladder: A reliable leaderboard for machine learning competitions
Blum, A.; and Hardt, M. 2015 · 2015
Cited alongside, same era.
Iyyer, M.; Wieting, J.; Gimpel, K.; and Zettlemoyer, L. 2018 · 2018
Later among the works it cites.
Show Your Work: Improved Reporting of Experimental Results
Dodge, J.; Gururangan, S.; Card, D.; Schwartz, R.; and Smith, N. A. 2019 · 2019
Later among the works it cites.
We need to talk about standard splits
Gorman, K.; and Bedrick, S. 2019 · 2019
Later among the works it cites.
Certified Robustness to Adversarial Word Substitutions
Jia, R.; Raghunathan, A.; Göksel, K.; and Liang, P. 2019 · 2019
Later among the works it cites.
Studying Summarization Evaluation Metrics in the Appropriate Scoring Range
Peyrard, M. 2019 · 2019
Later among the works it cites.
Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation Metrics
Mathur, N.; Baldwin, T.; and Cohn, T. 2020 · 2020
Later among the works it cites.
Do We Need to Create Big Datasets to Learn a Task?
Mishra, S.; and Sachdeva, B. S. 2020 · 2020
Later among the works it cites.
Beyond Accuracy: Behavioral Testing of NLP models with CheckList
Ribeiro, M. T.; Wu, T.; Guestrin, C.; and Singh, S. 2020 · 2020
Later among the works it cites.
BLEURT: Learning Robust Metrics for Text Generation
Sellam, T.; Das, D.; and Parikh, A. P. 2020 · 2020
Later among the works it cites.
Curriculum Learning for Natural Language Understanding
Xu, B.; Zhang, L.; Mao, Z.; Wang, Q.; Xie, H.; and Zhang, Y. 2020 · 2020
Later among the works it cites.