Fetching the paper…
Reading the bibliography…
Models that surpass human performance on several popular benchmarks display significant degradation in performance on exposure to Out of Distribution (OOD) data.
Distributional structure
Harris, Z. S · 1954
Earlier work this paper cites.
Convolutional networks for images, speech, and time series
LeCun, Y., Bengio, Y., et al · 1995
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Dataset shift in machine learning
Quionero-Candela, J., Sugiyama, M., Schwaighofer, A., and Lawrence, N. D · 2009
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Unbiased look at dataset bias
Torralba, A. and Efros, A. A · 2011
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C. D · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Bowman, S. R., Angeli, G., Potts, C., and Manning, C. D · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al · 2015
Earlier work this paper cites.
Towards universal paraphrastic sentence embeddings
Wieting, J., Bansal, M., Gimpel, K., and Livescu, K · 2015
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Adversarial examples for evaluating reading comprehension systems
Jia, R. and Liang, P · 2017
Cited alongside, same era.
The effect of different writing tasks on linguistic style: A case study of the roc story cloze task
Schwartz, R., Sap, M., Konstas, I., Zilles, L., Choi, Y., and Smith, N. A · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Robust physical-world attacks on deep learning visual classification
Eykholt, K., Evtimov, I., Fernandes, E., Li, B., Rahmati, A., Xiao, C., Prakash, A., Kohno, T., and Song, D · 2018
Cited alongside, same era.
Hendrycks, D., Zhao, K., Basart, S., Steinhardt, J., and Song, D · 2019
Later among the works it cites.
Learning the difference that makes a difference with counterfactually-augmented data
Kaushik, D., Hovy, E., and Lipton, Z. C · 2019
Later among the works it cites.
Repair: Removing representation bias by dataset resampling
Li, Y. and Vasconcelos, N · 2019
Later among the works it cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Later among the works it cites.
simple but effective techniques to reduce biases
Mahabadi, R. K. and Henderson, J · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gururangan, S., Swayamdipta, S., Levy, O., Schwartz, R., Bowman, S. R., and Smith, N. A · 2018
Cited alongside, same era.
How much reading does reading comprehension require? a critical investigation of popular benchmarks
Kaushik, D. and Lipton, Z. C · 2018
Cited alongside, same era.
Resound: Towards action recognition without representation bias
Li, Y., Li, Y., and Vasconcelos, N · 2018
Cited alongside, same era.
Hypothesis only baselines in natural language inference
Poliak, A., Naradowsky, J., Haldar, A., Rudinger, R., and Van Durme, B · 2018
Cited alongside, same era.
Performance impact caused by hidden bias of training data for recognizing textual entailment
Tsuchiya, M · 2018
Cited alongside, same era.
Wang, T., Zhu, J.-Y., Torralba, A., and Efros, A. A · 2018
Cited alongside, same era.
Don’t take the easy way out: Ensemble based methods for avoiding known dataset biases
Clark, C., Yatskar, M., and Zettlemoyer, L · 2019
Cited alongside, same era.
We need to talk about standard splits
Gorman, K. and Bedrick, S · 2019
Cited alongside, same era.
Later among the works it cites.
Adversarial nli: A new benchmark for natural language understanding
Nie, Y., Williams, A., Dinan, E., Bansal, M., Weston, J., and Kiela, D · 2019
Later among the works it cites.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2019
Later among the works it cites.
Multiqa: An empirical investigation of generalization and transfer in reading comprehension
Talmor, A. and Berant, J · 2019
Later among the works it cites.
Investigating biases in textual entailment datasets
Tan, S., Shen, Y., Huang, C.-w., and Courville, A · 2019
Later among the works it cites.
Adversarial filters of dataset biases
Bras, R. L., Swayamdipta, S., Bhagavatula, C., Zellers, R., Peters, M. E., Sabharwal, A., and Choi, Y · 2020
Closest in time.
Evaluating nlp models via contrast sets
Gardner, M., Artzi, Y., Basmova, V., Berant, J., Bogin, B., Chen, S., Dasigi, P., Dua, D., Elazar, Y., Gottumukkala, A., et al · 2020
Closest in time.
Pretrained transformers improve out-of-distribution robustness
Hendrycks, D., Liu, X., Wallace, E., Dziedzic, A., Krishnan, R., and Song, D · 2020
Closest in time.
Dqi: Measuring data quality in nlp
Mishra, S., Arunkumar, A., Sachdeva, B., Bryan, C., and Baral, C · 2020
Closest in time.
Train in germany, test in the usa: Making 3d object detectors generalize
Wang, Y., Chen, X., You, Y., Li, L. E., Hariharan, B., Campbell, M., Weinberger, K. Q., and Chao, W.-L · 2020
Closest in time.