The limits of machine intelligence
Henry Shevlin, Karina Vold, Matthew Crosby, and Marta Halina · 2019
Later among the works it cites.
Energy and policy considerations for deep learning in NLP
Emma Strubell, Ananya Ganesh, and Andrew McCallum · 2019
Later among the works it cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2019
Later among the works it cites.
Climbing towards NLU: On meaning, form, and understanding in the age of data
Emily M. Bender and Alexander Koller · 2020
Later among the works it cites.
Experience grounds language
Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, Nicolas Pinto, and Joseph Turian · 2020
Later among the works it cites.
Value-laden disciplinary shifts in machine learning
Ravit Dotan and Smitha Milli · 2020
Later among the works it cites.
General purpose text embeddings from pre-trained language models for scalable inference
Jingfei Du, Myle Ott, Haoran Li, Xing Zhou, and Veselin Stoyanov · 2020
Later among the works it cites.
Utility is in the eye of the user: A critique of NLP leaderboard design
Kawin Ethayarajh and Dan Jurafsky · 2020
Later among the works it cites.
What BERT is not: Lessons from a new suite of psycholinguistic diagnostics for language models
Allyson Ettinger · 2020
Later among the works it cites.
Race and gender
Timnit Gebru · 2020
Later among the works it cites.
Datasheets for datasets, 2020
Original
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford · 2020
Later among the works it cites.
Towards the systematic reporting of the energy and carbon footprints of machine learning
Peter Henderson, Jieru Hu, Joshua Romoff, Emma Brunskill, Dan Jurafsky, and Joelle Pineau · 2020
Later among the works it cites.
It’s not a non-issue: Negation as a source of error in machine translation
Md Mosharaf Hossain, Antonios Anastasopoulos, Eduardo Blanco, and Alexis Palmer · 2020
Later among the works it cites.
Social biases in NLP models as barriers for persons with disabilities
Ben Hutchinson, Vinodkumar Prabhakaran, Emily Denton, Kellie Webster, Yu Zhong, and Stephen Denuyl · 2020
Later among the works it cites.
Lessons from archives: Strategies for collecting sociocultural data in machine learning
Eun Seo Jo and Timnit Gebru · 2020
Later among the works it cites.
Deep learning for generic object detection: A survey
Li Liu, Wanli Ouyang, Xiaogang Wang, Paul Fieguth, Jie Chen, Xinwang Liu, and Matti Pietikäinen · 2020
Later among the works it cites.
Data and its (dis)contents: A survey of dataset development and use in machine learning research
Original
Amandalynne Paullada, Inioluwa Deborah Raji, Emily M Bender, Emily Denton, and Alex Hanna · 2020
Later among the works it cites.
Pre-trained models for natural language processing: A survey
XiPeng Qiu, TianXiang Sun, YiGe Xu, YiGe Shao, Ning Dai, and XuanJing Huang · 2020
Later among the works it cites.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh · 2020
Later among the works it cites.
Green AI
Roy Schwartz, Jesse Dodge, Noah A. Smith, and Oren Etzioni · 2020
Later among the works it cites.
Disembodied machine learning: On the illusion of objectivity in NLP, 2020
Zeerak Waseem, Smarika Lulz, Joachim Bingel, and Isabelle Augenstein · 2020
Later among the works it cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Closest in time.
Large image datasets: A pyrrhic win for computer vision?
Abeba Birhane and Vinay Uday Prabhu · 2021
Closest in time.
On the opportunities and risks of foundation models
Original
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Closest in time.
What will it take to fix benchmarking in natural language understanding?
Samuel R. Bowman and George Dahl · 2021
Closest in time.
On the genealogy of machine learning datasets: A critical history of imagenet
Emily Denton, Alex Hanna, Razvan Amironesei, Andrew Smart, and Hilary Nicole · 2021
Closest in time.
Microsoft deberta surpasses human performance on the superglue benchmark, 2021
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen · 2021
Closest in time.
Measurement and fairness
Abigail Z. Jacobs and Hanna Wallach · 2021
Closest in time.
Why AI is harder than we think
Original
Melanie Mitchell · 2021
Closest in time.
“everyone wants to do the model work, not the data work”: Data cascades in high-stakes AI
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M Aroyo · 2021
Closest in time.
Do datasets have politics? Disciplinary values in computer vision dataset development
Morgan Klaus Scheuerman, Alex Hanna, and Emily Denton · 2021
Closest in time.
Targeting the benchmark: On methodology in current natural language processing research
David Schlangen · 2021
Closest in time.
Underreporting of errors in NLG output, and what to do about it
Emiel van Miltenburg, Miruna Clinciu, Ondřej Dušek, Dimitra Gkatzia, Stephanie Inglis, Leo Leppänen, Saad Mahamood, Emma Manning, Stephanie Schoch, Craig Thomson, and Luou Wen · 2021
Closest in time.
TuringAdvice: A generative and dynamic evaluation of language use
Rowan Zellers, Ari Holtzman, Elizabeth Clark, Lianhui Qin, Ali Farhadi, and Yejin Choi · 2021
Closest in time.