Fetching the paper…
Reading the bibliography…
Benchmarks such as GLUE have helped drive advances in NLP by incentivizing the creation of more accurate models.
Roy Schwartz, Jesse Dodge, Noah A. Smith, and Oren Etzioni. 2019 · 1907
Earlier work this paper cites.
Adversarial nli: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2019 · 1910
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Fairness without demographics in repeated loss minimization
Tatsunori Hashimoto, Megha Srivastava, Hongseok Namkoong, and Percy Liang. 2018 · 1938
Earlier work this paper cites.
Consumption theory in terms of revealed preference
Paul A Samuelson. 1948 · 1948
Earlier work this paper cites.
Overview of results of the MUC-6 evaluation
Beth Sundheim. 1995 · 1995
Earlier work this paper cites.
Justice as fairness: A restatement
John Rawls. 2001 · 2001
Earlier work this paper cites.
Electra: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. 2020 · 2003
Earlier work this paper cites.
Is your classifier actually biased? measuring fairness under uncertainty with bernstein bounds
Kawin Ethayarajh. 2020 · 2004
Earlier work this paper cites.
Dynabert: Dynamic bert with adaptive width and depth
Lu Hou, Lifeng Shang, Xin Jiang, and Qun Liu. 2020 · 2004
Earlier work this paper cites.
Ladabert: Lightweight adaptation of bert through hybrid model compression
Yihuan Mao, Yujing Wang, Chufan Wu, Chen Zhang, Yang Wang, Yaming Yang, Quanlu Zhang, Yunhai Tong, and Jing Bai. 2020 · 2004
Earlier work this paper cites.
The effect of natural distribution shift on question answering models
John Miller, Karl Krauth, Benjamin Recht, and Ludwig Schmidt. 2020 · 2004
Earlier work this paper cites.
Stereoset: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy. 2020 · 2004
Earlier work this paper cites.
Gobo: Quantizing attention-based nlp models for low latency and energy efficient inference
Ali Hadi Zadeh and Andreas Moshovos. 2020 · 2005
Earlier work this paper cites.
Part 5: Machine Translation Evaluation . Springer Science & Business Media
Bonnie Dorr. 2011 · 2011
Earlier work this paper cites.
Sem 2013 shared task: Semantic textual similarity, including a pilot on typed-similarity
Eneko Agirre, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, and Weiwei Guo. 2013 · 2013
Earlier work this paper cites.
Semeval-2014 task 10: Multilingual semantic textual similarity
Eneko Agirre, Carmen Banea, Claire Cardie, Daniel M Cer, Mona T Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Rada Mihalcea, German Rigau, and Janyce Wiebe. 2014 · 2014
Earlier work this paper cites.
Semeval-2015 task 2: Semantic textual similarity, English, Spanish and pilot on interpretability
Eneko Agirre, Carmen Banea, Claire Cardie, Daniel M Cer, Mona T Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Inigo Lopez-Gazpio, Montse Maritxalar, Rada Mihalcea, et al. 2015 · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning. 2015 · 2015
Earlier work this paper cites.
The ladder: a reliable leaderboard for machine learning competitions
Moritz Hardt and Avrim Blum. 2015 · 2015
Cited alongside, same era.
Speed or accuracy? a study in evaluation of simultaneous speech translation
Takashi Mieno, Graham Neubig, Sakriani Sakti, Tomoki Toda, and Satoshi Nakamura. 2015 · 2015
Cited alongside, same era.
Equality of opportunity in supervised learning
Moritz Hardt, Eric Price, and Nati Srebro. 2016 · 2016
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
A simple but tough-to-beat baseline for sentence embeddings
Sanjeev Arora, Yingyu Liang, and Tengyu Ma. 2019 · 2017
Cited alongside, same era.
Fairness in machine learning
Solon Barocas, Moritz Hardt, and Arvind Narayanan. 2017 · 2017
Cited alongside, same era.
Rotate king to get queen: Word relationships as orthogonal transformations in embedding space
Kawin Ethayarajh. 2019 · 2019
Later among the works it cites.
Certified robustness to adversarial word substitutions
Robin Jia, Aditi Raghunathan, Kerem Göksel, and Percy Liang. 2019 · 2019
Later among the works it cites.
Black is to criminal as caucasian is to police: Detecting and removing multiclass bias in word embeddings
Thomas Manzini, Lim Yao Chong, Alan W Black, and Yulia Tsvetkov. 2019 · 2019
Later among the works it cites.
Model cards for model reporting
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019 · 2019
Later among the works it cites.
Distributionally robust language modeling
Yonatan Oren, Shiori Sagawa, Tatsunori Hashimoto, and Percy Liang. 2019 · 2019
Later among the works it cites.
Energy and policy considerations for deep learning in nlp
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Climbing a shaky ladder: Better adaptive risk estimation
Moritz Hardt. 2017 · 2017
Cited alongside, same era.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M Bender and Batya Friedman. 2018 · 2018
Cited alongside, same era.
Questionable Answers in Question Answering Research: Reproducibility and Variability of Published Results
Matt Crane. 2018 · 2018
Cited alongside, same era.
Unsupervised random walk sentence embeddings: A strong but simple baseline
Kawin Ethayarajh. 2018 · 2018
Cited alongside, same era.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumeé III, and Kate Crawford. 2018 · 2018
Cited alongside, same era.
The natural language decathlon: Multitask learning as question answering
Bryan McCann, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher. 2018 · 2018
Cited alongside, same era.
Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019 · 2019
Later among the works it cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019 · 2019
Later among the works it cites.
Language (technology) is power: A critical survey of ”bias” in nlp
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020 · 2020
Closest in time.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Closest in time.
Interpretable multi-dataset evaluation for named entity recognition
Jinlan Fu, Pengfei Liu, and Graham Neubig. 2020 · 2020
Closest in time.
Findings of the fourth workshop on neural generation and translation
Kenneth Heafield, Hiroaki Hayashi, Yusuke Oda, Ioannis Konstas, Andrew Finch, Graham Neubig, Xian Li, and Alexandra Birch. 2020 · 2020
Closest in time.
How can we accelerate progress towards human-like linguistic generalization?
Tal Linzen. 2020 · 2020
Closest in time.
Principles of economics
N Gregory Mankiw. 2020 · 2020
Closest in time.
Neurips 2020 efficientqa competition: Systems, analyses and lessons learned
Sewon Min, Jordan Boyd-Graber, Chris Alberti, Danqi Chen, Eunsol Choi, Michael Collins, Kelvin Guu, Hannaneh Hajishirzi, Kenton Lee, Jennimaria Palomaki, et al · 2020
Closest in time.
How the transformers broke NLP leaderboards
Anna Rogers. 2019 · 2020
Closest in time.
Peer review in NLP: reject-if-not-SOTA
Anna Rogers. 2020 · 2020
Closest in time.
Adversarial attacks on deep-learning models in natural language processing: A survey
Wei Emma Zhang, Quan Z Sheng, Ahoud Alhazmi, and Chenliang Li. 2020 · 2020
Closest in time.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2031
Closest in time.