Fetching the paper…
Reading the bibliography…
Recently, NLP models have achieved remarkable progress across a variety of tasks; however, they have also been criticized for being not robust.
The intraclass correlation coefficient as a measure of reliability
John J Bartko. 1966 · 1966
Earlier work this paper cites.
Wordnet: A lexical database for english
George A. Miller. 1995 · 1995
Earlier work this paper cites.
Wt5?! training text-to-text models to explain their predictions
Sharan Narang, Colin Raffel, Katherine Lee, Adam Roberts, Noah Fiedel, and Karishma Malkan. 2020 · 2004
Earlier work this paper cites.
Domain adaptation with structural correspondence learning
John Blitzer, Ryan McDonald, and Fernando Pereira. 2006 · 2006
Earlier work this paper cites.
Dbpedia: A nucleus for a web of open data
Sören Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak, and Zachary Ives. 2007 · 2007
Earlier work this paper cites.
Biographies, Bollywood, boom-boxes and blenders: Domain adaptation for sentiment classification
John Blitzer, Mark Dredze, and Fernando Pereira. 2007 · 2007
Earlier work this paper cites.
A survey on transfer learning
Sinno Jialin Pan and Qiang Yang. 2010 · 2010
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Yelp dataset challenge: Review rating prediction
Nabiha Asghar. 2016 · 2016
Earlier work this paper cites.
Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering
Ruining He and Julian McAuley. 2016 · 2016
Earlier work this paper cites.
Counter-fitting word vectors to linguistic constraints
Nikola Mrkšić, Diarmuid Ó Séaghdha, Blaise Thomson, Milica Gašić, Lina Rojas-Barahona, Pei-Hao Su, David Vandyke, Tsung-Hsien Wen, and Steve Young. 2016 · 2016
Earlier work this paper cites.
"why should I trust you?": Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Earlier work this paper cites.
Generating natural language adversarial examples
Moustafa Alzantot, Yash Sharma, Ahmed Elgohary, Bo-Jhang Ho, Mani Srivastava, and Kai-Wei Chang. 2018 · 2018
Earlier work this paper cites.
Fairwashing: the risk of rationalization
Ulrich Aïvodji, Hiromi Arai, Olivier Fortineau, Sébastien Gambs, Satoshi Hara, and Alain Tapp. 2019 · 2019
Earlier work this paper cites.
Don’t take the easy way out: Ensemble based methods for avoiding known dataset biases
Christopher Clark, Mark Yatskar, and Luke Zettlemoyer. 2019a · 2019
Cited alongside, same era.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019b · 2019
Cited alongside, same era.
Bias in bios
Maria De-Arteaga, Alexey Romanov, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Unlearn dataset bias in natural language inference by fitting the residual
He He, Sheng Zha, and Haohan Wang. 2019 · 2019
Cited alongside, same era.
Revealing the dark secrets of BERT
Automatic shortcut removal for self-supervised representation learning
Matthias Minderer, Olivier Bachem, Neil Houlsby, and Michael Tschannen. 2020 · 2020
Later among the works it cites.
Evaluating robustness to input perturbations for neural machine translation
Xing Niu, Prashant Mathur, Georgiana Dinu, and Yaser Al-Onaizan. 2020 · 2020
Later among the works it cites.
Learning to deceive with attention-based explanations
Danish Pruthi, Mansi Gupta, Bhuwan Dhingra, Graham Neubig, and Zachary C. Lipton. 2020 · 2020
Later among the works it cites.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Later among the works it cites.
An investigation of why overparameterization exacerbates spurious correlations
Shiori Sagawa, Aditi Raghunathan, Pang Wei Koh, and Percy Liang. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 2019
Cited alongside, same era.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 2019
Cited alongside, same era.
Robustness may be at odds with accuracy
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner, and Aleksander Madry. 2019 · 2019
Cited alongside, same era.
Adversarial domain adaptation for machine reading comprehension
Huazheng Wang, Zhe Gan, Xiaodong Liu, Jingjing Liu, Jianfeng Gao, and Hongning Wang. 2019 · 2019
Cited alongside, same era.
Adversarial filters of dataset biases
Ronan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers, Matthew Peters, Ashish Sabharwal, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Generating hierarchical explanations on text classification via feature interaction detection
Hanjie Chen, Guangtao Zheng, and Yangfeng Ji. 2020 · 2020
Cited alongside, same era.
Learning to model and ignore dataset bias with mixed capacity ensembles
Christopher Clark, Mark Yatskar, and Luke Zettlemoyer. 2020 · 2020
Cited alongside, same era.
Robustness to spurious correlations via human annotations
Megha Srivastava, Tatsunori Hashimoto, and Percy Liang. 2020 · 2020
Later among the works it cites.
An empirical study on robustness to spurious correlations using pre-trained language models
Lifu Tu, Garima Lalwani, Spandana Gella, and He He. 2020 · 2020
Later among the works it cites.
Identifying spurious correlations for robust text classification
Zhao Wang and Aron Culotta. 2020a · 2020
Later among the works it cites.
Towards robustifying NLI models against lexical dataset biases
Xiang Zhou and Mohit Bansal. 2020 · 2020
Later among the works it cites.
Towards interpreting and mitigating shortcut learning behavior of NLU models
Mengnan Du, Varun Manjunatha, Rajiv Jain, Ruchi Deshpande, Franck Dernoncourt, Jiuxiang Gu, Tong Sun, and Xia Hu. 2021 · 2021
Closest in time.
Self-attention attribution: Interpreting information interactions inside transformer
Yaru Hao, Li Dong, Furu Wei, and Ke Xu. 2020 · 2021
Closest in time.
Contrastive explanations for model interpretability
Alon Jacovi, Swabha Swayamdipta, Shauli Ravfogel, Yanai Elazar, Yejin Choi, and Yoav Goldberg. 2021 · 2021
Closest in time.
Hiddencut: Simple data augmentation for natural language understanding with better generalizability
Chen Jiaao, Shen Dinghan, Chen Weizhu, and Yang Diyi. 2021 · 2021
Closest in time.
Removing spurious features can hurt accuracy and affect groups disproportionately
Fereshte Khani and Percy Liang. 2021 · 2021
Closest in time.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2031
Closest in time.