Fetching the paper…
Reading the bibliography…
We examine multi-task benchmarks in machine learning through the lens of social choice theory.
A difficulty in the concept of social welfare
Kenneth J Arrow · 1950
Earlier work this paper cites.
Social choice and individual values
Kenneth J. Arrow · 1951
Earlier work this paper cites.
Social choice theory: An introduction
Jerry S Kelly · 1988
Earlier work this paper cites.
The historical perspective of the problem of interpersonal comparisons of utility
Stavros A. Drakopoulos · 1989
Earlier work this paper cites.
Mathematics without numbers
Stewart Shapiro and Geoffrey Hellman · 1993
Earlier work this paper cites.
Social choice and the mathematics of manipulation
Alan D Taylor · 2005
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, K. Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron C. Courville · 2013
Earlier work this paper cites.
Preserving statistical validity in adaptive data analysis
Cynthia Dwork, Vitaly Feldman, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Aaron Roth · 2014
Earlier work this paper cites.
The ladder: A reliable leaderboard for machine learning competitions
Avrim Blum and Moritz Hardt · 2015
Earlier work this paper cites.
The reusable holdout: Preserving validity in adaptive data analysis
Cynthia Dwork, Vitaly Feldman, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Aaron Roth · 2015
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Shane Gu, and Ben Poole · 2016
Earlier work this paper cites.
Torchvision: Pytorch’s computer vision library
TorchVision maintainers and contributors · 2016
Earlier work this paper cites.
Learning to score system summaries for better content selection evaluation
Maxime Peyrard, Teresa Botschen, and Iryna Gurevych · 2017
Earlier work this paper cites.
Collective choice and social welfare
Amartya Sen · 2017
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2018
Earlier work this paper cites.
The advantages of multiple classes for reducing overfitting from test set reuse
Vitaly Feldman, Roy Frostig, and Moritz Hardt · 2019
Earlier work this paper cites.
Model similarity mitigates test set overuse
Horia Mania, John Miller, Ludwig Schmidt, Moritz Hardt, and Benjamin Recht · 2019
Earlier work this paper cites.
Recent advances in natural language inference: A survey of benchmarks, resources, and approaches
Shane Storks, Qiaozi Gao, and Joyce Yue Chai · 2019
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Cited alongside, same era.
A large-scale study of representation learning with the visual task adaptation benchmark
Xiaohua Zhai, Joan Puigcerver, Alexander Kolesnikov, Pierre Ruyssen, Carlos Riquelme, Mario Lucic, Josip Djolonga, André Susano Pinto, Maxim Neumann, Alexey Dosovitskiy, Lucas Beyer, Olivier Bachem, Michael Tschannen, Marcin Michalski, Olivier Bousquet, Sylvain Gelly, and Neil Houlsby · 2019
Cited alongside, same era.
Machine learning testing: Survey, landscapes and horizons
J Zhang, Mark Harman, Lei Ma, and Yang Liu · 2019
Cited alongside, same era.
Utility is in the eye of the user: A critique of nlp leaderboard design
Kawin Ethayarajh and Dan Jurafsky · 2020
Cited alongside, same era.
Data contamination: From memorization to exploitation
Inbal Magar and Roy Schwartz · 2022
Later among the works it cites.
Mteb: Massive text embedding benchmark
Niklas Muennighoff, Nouamane Tazi, Loïc Magne, and Nils Reimers · 2022
Later among the works it cites.
Mapping global dynamics of benchmark creation and saturation in artificial intelligence
Simon Ott, Adriano Barbosa-Silva, Kathrin Blagec, Janina Brauner, and Matthias Samwald · 2022
Later among the works it cites.
Vote’n’rank: Revision of benchmarking with social choice theory
Mark Rofin, Vladislav Mikhailov, Mikhail Florinskiy, Andrey Kravchenko, E. Tutubalina, Tatiana Shavrina, Daniel Karabekyan, and E. Artemova · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Xiaodong Song, and Jacob Steinhardt · 2020
Cited alongside, same era.
From imagenet to image classification: Contextualizing progress on benchmarks
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Andrew Ilyas, and Aleksander Madry · 2020
Cited alongside, same era.
Rip van winkle’s razor: A simple estimate of overfit to test data
Sanjeev Arora and Yi Zhang · 2021
Cited alongside, same era.
What will it take to fix benchmarking in natural language understanding?
Samuel R. Bowman and George E. Dahl · 2021
Cited alongside, same era.
Infolm: A new metric to evaluate summarization & data2text generation
Pierre Colombo, Chloe Clave, and Pablo Piantanida · 2021
Cited alongside, same era.
Mostafa Dehghani, Yi Tay, Alexey A. Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, and Oriol Vinyals · 2021
Cited alongside, same era.
A framework for few-shot language model evaluation, September 2021
Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou · 2021
Cited alongside, same era.
Dynabench: Rethinking benchmarking in nlp
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Talat, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams · 2021
Cited alongside, same era.
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2022
Later among the works it cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
Mirac Suzgun, Nathan Scales, Nathanael Scharli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed Huai hsin Chi, Denny Zhou, and Jason Wei · 2022
Later among the works it cites.
Open llm leaderboard
Edward Beeching, Clémentine Fourrier, Nathan Habib, Sheon Han, Nathan Lambert, Nazneen Rajani, Omar Sanseviero, Lewis Tunstall, and Thomas Wolf · 2023
Later among the works it cites.
Elo uncovered: Robustness and best practices in language model evaluation
Meriem Boubdir, Edward Kim, Beyza Hilal Ermiş, Sara Hooker, and Marzieh Fadaee · 2023
Later among the works it cites.
Data science at the singularity
David Donoho · 2023
Later among the works it cites.
Towards more robust nlp system evaluation: Handling missing scores in benchmarks
Anas Himmi, Ekhine Irurozki, Nathan Noiry, Stéphan Clémençon, and Pierre Colombo · 2023
Later among the works it cites.
Holistic evaluation of text-to-image models
Tony Lee, Michihiro Yasunaga, Chenlin Meng, Yifan Mai, Joon Sung Park, Agrim Gupta, Yunzhi Zhang, Deepak Narayanan, Hannah Teufel, Marco Bellagente, Minguk Kang, Taesung Park, Jure Leskovec, Jun-Yan Zhu, Fei-Fei Li, Jiajun Wu, Stefano Ermon, and Percy Liang · 2023
Later among the works it cites.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Data contamination through the lens of time
Manley Roberts, Himanshu Thakur, Christine Herlihy, Colin White, and Samuel Dooley · 2023
Later among the works it cites.
A theory of dynamic benchmarks
Ali Shirali, Rediet Abebe, and Moritz Hardt · 2023
Later among the works it cites.
Gemini: A family of highly capable multimodal models
Gemini Team · 2023
Later among the works it cites.
Llmeval: A preliminary study on how to evaluate large language models
Yue Zhang, Ming Zhang, Haipeng Yuan, Shichun Liu, Yongyao Shi, Tao Gui, Qi Zhang, and Xuanjing Huang · 2023
Later among the works it cites.
When benchmarks are targets: Revealing the sensitivity of large language model leaderboards
Norah Alzahrani, Hisham Abdullah Alyahya, Sultan Yazeed Alnumay, Shaykhah Alsubaie, Yusef Almushaykeh, Faisal Mirza, Nouf Alotaibi, Nora Altwairesh, Areeb Alowisheq, Saiful Bari, and Haidar Khan · 2024
Closest in time.