Fetching the paper…
Reading the bibliography…
NLU models often exploit biases to achieve high dataset-specific performance without properly learning the intended task.
Adversarial nli: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2019b · 1910
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2019 · 1910
Earlier work this paper cites.
R Thomas McCoy, Junghyun Min, and Tal Linzen. 2019a · 1911
Earlier work this paper cites.
Robust natural language inference models with example forgetting
Yadollah Yaghoobzadeh, Remi Tachet, Timothy J Hazen, and Alessandro Sordoni. 2019 · 1911
Earlier work this paper cites.
Evaluating NLP models via contrast sets
Matt Gardner, Yoav Artzi, Victoria Basmova, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, et al. 2020 · 2004
Earlier work this paper cites.
The curse of performance instability in analysis datasets: Consequences, source, and suggestions
Xiang Zhou, Yixin Nie, Hao Tan, and Mohit Bansal. 2020 · 2004
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini. 2005 · 2005
Earlier work this paper cites.
On the value of out-of-distribution testing: An example of goodhart’s law
Damien Teney, Kushal Kafle, Robik Shrestha, Ehsan Abbasnejad, Christopher Kanan, and Anton van den Hengel. 2020 · 2005
Earlier work this paper cites.
The second pascal recognising textual entailment challenge
Roy Bar-Haim, Ido Dagan, Bill Dolan, Lisa Ferro, and Danilo Giampiccolo. 2006 · 2006
Earlier work this paper cites.
On the stability of fine-tuning bert: Misconceptions, explanations, and strong baselines
Marius Mosbach, Maksym Andriushchenko, and Dietrich Klakow. 2020 · 2006
Earlier work this paper cites.
Effectively using syntax for recognizing false entailment
Rion Snow, Lucy Vanderwende, and Arul Menezes. 2006 · 2006
Earlier work this paper cites.
What syntax can contribute in the entailment task
Lucy Vanderwende and William B. Dolan. 2006 · 2006
Earlier work this paper cites.
Revisiting few-sample bert fine-tuning
Tianyi Zhang, Felix Wu, Arzoo Katiyar, Kilian Q Weinberger, and Yoav Artzi. 2020 · 2006
Earlier work this paper cites.
The third PASCAL recognizing textual entailment challenge
Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and Bill Dolan. 2007 · 2007
Earlier work this paper cites.
Early-learning regularization prevents memorization of noisy labels
Sheng Liu, Jonathan Niles-Weed, Narges Razavian, and Carlos Fernandez-Granda. 2020 · 2007
Earlier work this paper cites.
A SICK cure for the evaluation of compositional distributional semantic models
Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, and Roberto Zamparelli. 2014 · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Earlier work this paper cites.
A closer look at memorization in deep networks
Devansh Arpit, Stanisław Jastrzundefinedbski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S. Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, and Simon Lacoste-Julien. 2017 · 2017
Cited alongside, same era.
Pay attention to the ending:strong neural baselines for the ROC story cloze task
Zheng Cai, Lifu Tu, and Kevin Gimpel. 2017 · 2017
Cited alongside, same era.
The effect of different writing tasks on linguistic style: A case study of the ROC story cloze task
Roy Schwartz, Maarten Sap, Ioannis Konstas, Leila Zilles, Yejin Choi, and Noah A. Smith. 2017 · 2017
Cited alongside, same era.
Evaluating compositionality in sentence embeddings
Ishita Dasgupta, Demi Guo, Andreas Stuhlmüller, Samuel J Gershman, and Noah D. Goodman. 2018 · 2018
Cited alongside, same era.
Adversarial removal of demographic attributes from text data
Yanai Elazar and Yoav Goldberg. 2018 · 2018
Cited alongside, same era.
SWAG: A large-scale adversarial dataset for grounded commonsense inference
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018 · 2018
Later among the works it cites.
Don’t take the easy way out: Ensemble based methods for avoiding known dataset biases
Christopher Clark, Mark Yatskar, and Luke Zettlemoyer. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Unlearn dataset bias in natural language inference by fitting the residual
He He, Sheng Zha, and Haohan Wang. 2019 · 2019
Later among the works it cites.
Learning the difference that makes a difference with counterfactually-augmented data
Divyansh Kaushik, Eduard Hovy, and Zachary Lipton. 2020 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Born-again neural networks
Tommaso Furlanello, Zachary Chase Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar. 2018 · 2018
Cited alongside, same era.
Breaking NLI systems with sentences that require simple lexical inferences
Max Glockner, Vered Shwartz, and Yoav Goldberg. 2018 · 2018
Cited alongside, same era.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018 · 2018
Cited alongside, same era.
How much reading does reading comprehension require? a critical investigation of popular benchmarks
Divyansh Kaushik and Zachary C. Lipton. 2018 · 2018
Cited alongside, same era.
Scitail: A textual entailment dataset from science question answering
Tushar Khot, Ashish Sabharwal, and Peter Clark. 2018 · 2018
Cited alongside, same era.
Stress test evaluation for natural language inference
Aakanksha Naik, Abhilasha Ravichander, Norman Sadeh, Carolyn Rose, and Graham Neubig. 2018 · 2018
Cited alongside, same era.
Hypothesis only baselines in natural language inference
Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018 · 2018
Cited alongside, same era.
Timothy Niven and Hung-Yu Kao. 2019 · 2019
Later among the works it cites.
Towards debiasing fact verification models
Tal Schuster, Darsh Shah, Yun Jie Serene Yeo, Daniel Roberto Filizzola Ortiz, Enrico Santus, and Regina Barzilay. 2019 · 2019
Later among the works it cites.
HellaSwag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Later among the works it cites.
PAWS: Paraphrase adversaries from word scrambling
Yuan Zhang, Jason Baldridge, and Luheng He. 2019 · 2019
Later among the works it cites.
Adversarial filters of dataset biases
Ronan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers, Matthew E. Peters, Ashish Sabharwal, and Yejin Choi. 2020 · 2020
Closest in time.
End-to-end bias mitigation by modelling biases in corpora
Rabeeh Mahabadi, Yonatan Belinkov, and James Henderson. 2020 · 2020
Closest in time.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020 · 2020
Closest in time.
Predictive biases in natural language processing models: A conceptual framework and overview
Deven Santosh Shah, H. Andrew Schwartz, and Dirk Hovy. 2020 · 2020
Closest in time.
An empirical study on robustness to spurious correlations using pre-trained language models
Lifu Tu, Garima Lalwani, Spandana Gella, and He He. 2020 · 2020
Closest in time.
Mind the trade-off: Debiasing NLU models without degrading the in-distribution performance
Prasetya Ajie Utama, Nafise Sadat Moosavi, and Iryna Gurevych. 2020 · 2020
Closest in time.
Improving QA generalization by concurrent modeling of multiple biases
Mingzhu Wu, Nafise Sadat Moosavi, Andreas Rücklé, and Iryna Gurevych. 2020 · 2020
Closest in time.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2031
Closest in time.