Crosscheck: Rapid, reproducible, and interpretable model evaluation
Original
Dustin Arendt, Zhuanyi Huang, Prasha Shrestha, E. Ayton, Maria Glenski, and Svitlana Volkova · 2020
Later among the works it cites.
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Later among the works it cites.
Summeval: Re-evaluating summarization evaluation
Original
Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev · 2020
Later among the works it cites.
Empirical evaluation of pretraining strategies for supervised entity linking
Original
Thibault Févry, Nicholas FitzGerald, Livio Baldini Soares, and T. Kwiatkowski · 2020
Later among the works it cites.
Evaluating nlp models via contrast sets
Original
Matt Gardner, Yoav Artzi, Victoria Basmova, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, et al · 2020
Later among the works it cites.
Bae: Bert-based adversarial examples for text classification
Original
Siddhant Garg and Goutham Ramakrishnan · 2020
Later among the works it cites.
Twitter is investigating after anecdotal data suggested its picture-cropping tool favors white faces, 2020
Isobel Asher Hamilton · 2020
Later among the works it cites.
The many faces of robustness: A critical analysis of out-of-distribution generalization
Original
Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, et al · 2020
Later among the works it cites.
spaCy: Industrial-strength Natural Language Processing in Python, 2020
Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd · 2020
Later among the works it cites.
Are natural language inference models imppressive? learning implicature and presupposition
Original
Paloma Jeretic, Alex Warstadt, Suvrat Bhooshan, and Adina Williams · 2020
Later among the works it cites.
Google apologizes after its Vision AI produced racist results, 2020
Nicolas Kayser-Bril · 2020
Later among the works it cites.
Rethinking AI Benchmarking, 2020
Douwe Kiela et al · 2020
Later among the works it cites.
Wilds: A benchmark of in-the-wild distribution shifts, 2020
Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Sara Beery, Jure Leskovec, Anshul Kundaje, Emma Pierson, Sergey Levine, Chelsea Finn, and Percy Liang · 2020
Later among the works it cites.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Original
M. Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, A. Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer · 2020
Later among the works it cites.
The effect of natural distribution shift on question answering models
Original
J. Miller, Karl Krauth, B. Recht, and L. Schmidt · 2020
Later among the works it cites.
Textattack: A framework for adversarial attacks in natural language processing
Original
John X Morris, Eli Lifland, Jin Yong Yoo, and Yanjun Qi · 2020
Later among the works it cites.
Explaining machine learning classifiers through diverse counterfactual explanations
Ramaravind K Mothilal, Amit Sharma, and Chenhao Tan · 2020
Later among the works it cites.
Adversarial nli: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela · 2020
Later among the works it cites.
Bootleg: Chasing the tail with self-supervised named entity disambiguation
Original
L. Orr, Megan Leszczynski, Simran Arora, Sen Wu, N. Guha, Xiao Ling, and C. Ré · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, M. Matena, Yanqi Zhou, W. Li, and Peter J. Liu · 2020
Later among the works it cites.
Beyond accuracy: Behavioral testing of nlp models with checklist
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh · 2020
Later among the works it cites.
Conjnli: Natural language inference over conjunctive sentences
Swarnadeep Saha, Yixin Nie, and Mohit Bansal · 2020
Later among the works it cites.
Data augmentation for discrimination prevention and bias disambiguation
Shubham Sharma, Yunfeng Zhang, Jesús M. Ríos Aliaga, Djallel Bouneffouf, Vinod Muthusamy, and Kush R. Varshney · 2020
Later among the works it cites.
Measuring robustness to natural distribution shifts in image classification
Original
Rohan Taori, Achal Dave, V. Shankar, N. Carlini, B. Recht, and L. Schmidt · 2020
Later among the works it cites.
The language interpretability tool: Extensible, interactive visualizations and analysis for nlp models
Original
Ian Tenney, James Wexler, Jasmijn Bastings, Tolga Bolukbasi, Andy Coenen, Sebastian Gehrmann, Ellen Jiang, Mahima Pushkarna, Carey Radebaugh, Emily Reif, et al · 2020
Later among the works it cites.
Rel: An entity linker standing on the shoulders of giants
Johannes M. van Hulst, F. Hasibi, K. Dercksen, K. Balog, and A. D. Vries · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2020
Later among the works it cites.
Do neural models learn systematicity of monotonicity inference in natural language?
Original
Hitomi Yanaka, Koji Mineshima, Daisuke Bekki, and Kentaro Inui · 2020
Later among the works it cites.
Openattack: An open-source textual adversarial attack toolkit
Original
Guoyang Zeng, Fanchao Qi, Qianrui Zhou, Tingji Zhang, Bairu Hou, Yuan Zang, Zhiyuan Liu, and Maosong Sun · 2020
Later among the works it cites.
Pegasus: Pre-training with extracted gap-sentences for abstractive summarization
Original
Jingqing Zhang, Y. Zhao, Mohammad Saleh, and Peter J. Liu · 2020
Later among the works it cites.