Fetching the paper…
Reading the bibliography…
Although measuring held-out accuracy has been the primary approach to evaluate generalization, it often overestimates the performance of NLP models, while alternative approaches for evaluating models either focus on individual tasks or on specific behaviors.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R’emi Louf, Morgan Funtowicz, and Jamie Brew. 2019 · 1910
Earlier work this paper cites.
Black-box Testing: Techniques for Functional Testing of Software and Systems
Boris Beizer. 1995 · 1995
Earlier work this paper cites.
Investigating statistical machine learning as a tool for software development
Kayur Patel, James Fogarty, James A Landay, and Beverly Harrison. 2008 · 2008
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Why should i trust you?: Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
A survey on metamorphic testing
Sergio Segura, Gordon Fraser, Ana B Sanchez, and Antonio Ruiz-Cortés. 2016 · 2016
Earlier work this paper cites.
Correlation-based intrinsic evaluation of word vector representations
Yulia Tsvetkov, Manaal Faruqui, and Chris Dyer. 2016 · 2016
Earlier work this paper cites.
Synthetic and natural noise both break neural machine translation
Yonatan Belinkov and Yonatan Bisk. 2018 · 2018
Earlier work this paper cites.
Adversarial example generation with syntactically controlled paraphrase networks
Mohit Iyyer, John Wieting, Kevin Gimpel, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Stress Test Evaluation for Natural Language Inference
Aakanksha Naik, Abhilasha Ravichander, Norman Sadeh, Carolyn Rose, and Graham Neubig. 2018 · 2018
Cited alongside, same era.
Know what you don’t know: Unanswerable questions for SQuAD
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Cited alongside, same era.
Semantically equivalent adversarial rules for debugging nlp models
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2018 · 2018
Cited alongside, same era.
What’s in your embedding, and how it predicts task performance
Anna Rogers, Shashwath Hosur Ananthakrishna, and Anna Rumshisky. 2018 · 2018
Cited alongside, same era.
Software engineering for machine learning: A case study
Saleema Amershi, Andrew Begel, Christian Bird, Rob DeLine, Harald Gall, Ece Kamar, Nachi Nagappan, Besmira Nushi, and Tom Zimmermann. 2019 · 2019
Probing what different nlp tasks teach machines about function word comprehension
Najoung Kim, Roma Patel, Adam Poliak, Patrick Xia, Alex Wang, Tom McCoy, Ian Tenney, Alexis Ross, Tal Linzen, Benjamin Van Durme, et al. 2019 · 2019
Later among the works it cites.
Perturbation sensitivity analysis to detect unintended model biases
Vinodkumar Prabhakaran, Ben Hutchinson, and Margaret Mitchell. 2019 · 2019
Later among the works it cites.
Do imagenet classifiers generalize to imagenet?
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. 2019 · 2019
Later among the works it cites.
Are red roses red? evaluating consistency of question-answering models
Marco Tulio Ribeiro, Carlos Guestrin, and Sameer Singh. 2019 · 2019
Later among the works it cites.
Models in the wild: On corruption robustness of neural nlp systems
Barbara Rychalska, Dominika Basaj, Alicja Gosiewska, and Przemysław Biecek. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Analysis methods in neural language processing: A survey
Yonatan Belinkov and James Glass. 2019 · 2019
Cited alongside, same era.
Are we modeling the task or the annotator? an investigation of annotator bias in natural language understanding datasets
Mor Geva, Yoav Goldberg, and Jonathan Berant. 2019 · 2019
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019a
Cited in the paper.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019b
Cited in the paper.
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Later among the works it cites.
Universal adversarial triggers for attacking and analyzing nlp
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019 · 2019
Later among the works it cites.
Errudite: Scalable, reproducible, and testable error analysis
Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel S Weld. 2019 · 2019
Later among the works it cites.