Fetching the paper…
Reading the bibliography…
With the advent of Transformers, large language models (LLMs) have saturated well-known NLP benchmarks and leaderboards with high aggregate performance.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
The spotlight: A general method for discovering systematic errors in deep learning models
Greg d’Eon, Jason d’Eon, James R Wright, and Kevin Leyton-Brown. 2022 · 1981
Earlier work this paper cites.
Latent dirichlet allocation
David M. Blei, Andrew Y. Ng, Michael I. Jordan, and John Lafferty. 2003 · 2003
Earlier work this paper cites.
Stability of k k -means clustering
Alexander Rakhlin and Andrea Caponnetto. 2006 · 2006
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Hidden factors and hidden topics: Understanding rating dimensions with review text
Julian McAuley and Jure Leskovec. 2013 · 2013
Earlier work this paper cites.
Character-level Convolutional Networks for Text Classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Amazon built an AI tool to hire people but had to shut it down because it was discriminating against women
Isobel Asher Hamilton. 2018 · 2018
Earlier work this paper cites.
Microsoft’s politically correct chatbot is even worse than its racist one
Chloe Rose Stuart-Ulin. 2018 · 2018
Earlier work this paper cites.
Facebook’s ad-serving algorithm discriminates by gender and race
Karen Hao. 2019 · 2019
Earlier work this paper cites.
The Apple Card Didn’t ’See’ Gender—and That’s the Problem
Will Knight. 2019 · 2019
Cited alongside, same era.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 2019
Cited alongside, same era.
Errudite: Scalable, reproducible, and testable error analysis
Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel S Weld. 2019 · 2019
Cited alongside, same era.
Twitter is investigating after anecdotal data suggested its picture-cropping tool favors white faces
Isobel Asher Hamilton. 2020 · 2020
Cited alongside, same era.
Learning the difference that makes a difference with counterfactually-augmented data
Divyansh Kaushik, Eduard Hovy, and Zachary Lipton. 2020 · 2020
Cited alongside, same era.
Google apologizes after its Vision AI produced racist results
Goodwill hunting: Analyzing and repurposing off-the-shelf named entity linking systems
Karan Goel, Laurel Orr, Nazneen Fatema Rajani, Jesse Vig, and Christopher Ré. 2021a · 2021
Later among the works it cites.
Dynabench: Rethinking benchmarking in NLP
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams. 2021 · 2021
Later among the works it cites.
Bigscience large open-science open-access multilingual language model
BigScience. 2022 · 2022
Closest in time.
Domino: Discovering systematic errors with cross-modal embeddings
Sabri Eyuboglu, Maya Varma, Khaled Saab, Jean-Benoit Delbrouck, Christopher Lee-Messer, Jared Dunnmon, James Zou, and Christopher Ré. 2022 · 2022
Closest in time.
Metashift: A dataset of datasets for evaluating contextual distribution shifts and training conflicts
Weixin Liang and James Zou. 2022 · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nicolas Kayser-Bril. 2020 · 2020
Cited alongside, same era.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 2020
Cited alongside, same era.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Cited alongside, same era.
Dataset cartography: Mapping and diagnosing datasets with training dynamics
Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Dall·e mini
Boris Dayma, Suraj Patil, Pedro Cuenca, Khalid Saifullah, Tanishq Abraham, Phuc Le Khac, Luke Melas, and Ritobrata Ghosh. 2021 · 2021
Cited alongside, same era.
Robustness gym: Unifying the nlp evaluation landscape
Karan Goel, Nazneen Rajani, Jesse Vig, Samson Tan, Jason Wu, Stephan Zheng, Caiming Xiong, Mohit Bansal, and Christopher Ré. 2021b
Cited in the paper.
Topic discovery via latent space clustering of pretrained language model representations
Yu Meng, Yunyi Zhang, Jiaxin Huang, Yu Zhang, and Jiawei Han. 2022 · 2022
Closest in time.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Closest in time.
iSEA: An interactive pipeline for semantic error analysis of NLP models
Jun Yuan, Jesse Vig, and Nazneen Rajani. 2022 · 2022
Closest in time.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2031
Closest in time.