Fetching the paper…
Reading the bibliography…
Much recent work in NLP has documented dataset artifacts, bias, and spurious correlations between input features and output labels.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, M. Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 1909
Earlier work this paper cites.
Teoria statistica delle classi e calcolo delle probabilita
C. E. Bonferroni. 1936 · 1936
Earlier work this paper cites.
The complexity of Boolean functions
Ingo Wegener. 1987 · 1987
Earlier work this paper cites.
Prepositional phrase attachment through a backed-off model
Michael Collins and James Brooks. 1995 · 1995
Earlier work this paper cites.
A measure for the complexity of boolean functions related to their implementation in neural networks
Leonardo Franco. 2001 · 2001
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Ido Dagan, Oren Glickman, and B. Magnini. 2005 · 2005
Earlier work this paper cites.
Dataset Shift in Machine Learning
Joaquin Quionero-Candela, Masashi Sugiyama, Anton Schwaighofer, and N. Lawrence. 2009 · 2009
Earlier work this paper cites.
Explaining the efficacy of counterfactually-augmented data
Divyansh Kaushik, Amrith Rajagopal Setlur, E. Hovy, and Zachary Chase Lipton. 2020 · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Analysis of boolean functions
Ryan O’Donnell. 2014 · 2014
Earlier work this paper cites.
A gold standard dependency corpus for English
Natalia Silveira, Timothy Dozat, Marie-Catherine de Marneffe, Samuel Bowman, Miriam Connor, John Bauer, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
A thorough examination of the CNN/Daily Mail reading comprehension task
Danqi Chen, Jason Bolton, and Christopher D. Manning. 2016 · 2016
Earlier work this paper cites.
The E2E dataset: New challenges for end-to-end generation
Jekaterina Novikova, Ondřej Dušek, and Verena Rieser. 2017 · 2017
Earlier work this paper cites.
How grammatical is character-level neural machine translation? Assessing MT quality with contrastive translation pairs
Rico Sennrich. 2017 · 2017
Earlier work this paper cites.
Foil it! Find One mismatch between image and language caption
Ravi Shekhar, Sandro Pezzelle, Yauhen Klimovich, Aurélie Herbelot, Moin Nabi, Enver Sangineto, and Raffaella Bernardi. 2017 · 2017
Cited alongside, same era.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Joy Buolamwini and Timnit Gebru. 2018 · 2018
Cited alongside, same era.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018 · 2018
Cited alongside, same era.
Hypothesis only baselines in natural language inference
Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018 · 2018
Cited alongside, same era.
Anchors: High-precision model-agnostic explanations
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2018 · 2018
Cited alongside, same era.
Gender bias in coreference resolution
Universal adversarial triggers for attacking and analyzing NLP
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019 · 2019
Later among the works it cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Later among the works it cites.
Learning to model and ignore dataset bias with mixed capacity ensembles
Christopher Clark, Mark Yatskar, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
Evaluating models’ local decision boundaries via contrast sets
Matt Gardner, Yoav Artzi, Victoria Basmov, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, Nitish Gupta, Hannaneh Hajishirzi, Gabriel Ilharco, Daniel Khashabi, Kevin Lin, Jiangming Liu, Nelson F. Liu, Phoebe Mulcaire, Qiang Ning, Sameer Singh, Noah A. Smith, Sanjay Subramanian, Reut Tsarfaty, Eric Wallace, Ally Zhang, and Ben Zhou. 2020 · 2020
Later among the works it cites.
Counterfactually-augmented SNLI training data does not yield better generalization than unaugmented data
William Huang, Haokun Liu, and Samuel R. Bowman. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018 · 2018
Cited alongside, same era.
What makes reading comprehension questions easier?
Saku Sugawara, Kentaro Inui, Satoshi Sekine, and Akiko Aizawa. 2018 · 2018
Cited alongside, same era.
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018 · 2018
Cited alongside, same era.
Don’t take the easy way out: Ensemble based methods for avoiding known dataset biases
Christopher Clark, Mark Yatskar, and Luke Zettlemoyer. 2019b · 2019
Cited alongside, same era.
Proceedings of the First Workshop on Gender Bias in Natural Language Processing . Association for Computational Linguistics, Florence, Italy
Marta R. Costa-jussà, Christian Hardmeier, Will Radford, and Kellie Webster, editors. 2019 · 2019
Cited alongside, same era.
Racial bias in hate speech and abusive language detection datasets
Thomas Davidson, Debasmita Bhattacharya, and Ingmar Weber. 2019 · 2019
Cited alongside, same era.
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
End-to-end bias mitigation by modelling biases in corpora
Rabeeh Karimi Mahabadi, Yonatan Belinkov, and James Henderson. 2020 · 2020
Later among the works it cites.
Adversarial filters of dataset biases
Ronan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers, Matthew E. Peters, Ashish Sabharwal, and Yejin Choi. 2020 · 2020
Later among the works it cites.
Textattack: A framework for adversarial attacks, data augmentation, and adversarial training in nlp
John Morris, Eli Lifland, Jin Yong Yoo, Jake Grigsby, Di Jin, and Yanjun Qi. 2020 · 2020
Later among the works it cites.
Improving compositional generalization in semantic parsing
Inbar Oren, Jonathan Herzig, Nitish Gupta, Matt Gardner, and Jonathan Berant. 2020 · 2020
Later among the works it cites.
Dataset cartography: Mapping and diagnosing datasets with training dynamics
Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, and Yejin Choi. 2020 · 2020
Later among the works it cites.
Learning from task descriptions
Orion Weller, Nicholas Lourie, Matt Gardner, and Matthew Peters. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020 · 2020
Later among the works it cites.
A too-good-to-be-true prior to reduce shortcut reliance
Nikolay Dagaev, Brett D. Roads, Xiaoliang Luo, Daniel N. Barry, Kaustubh R. Patil, and Bradley C. Love. 2021 · 2021
Closest in time.
Sensitivity as a complexity measure for sequence classification tasks
Michael Hahn, Dan Jurafsky, and Richard Futrell. 2021 · 2021
Closest in time.
Damien Teney, Ehsan Abbasnejad, Simon Lucey, and Anton van den Hengel. 2021 · 2021
Closest in time.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2031
Closest in time.