Fetching the paper…
Reading the bibliography…
Standardized benchmarks drive progress in machine learning.
The complexity of the pigeonhole principle
Miklós Ajtai · 1994
Earlier work this paper cites.
Computerized adaptive testing: Theory and practice
Wim J Van der Linden and Cees AW Glas · 2000
Earlier work this paper cites.
The basics of item response theory
Frank B Baker · 2001
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
How do humans teach: On curriculum learning and teaching dimension
Faisal Khan, Bilge Mutlu, and Jerry Zhu · 2011
Earlier work this paper cites.
Unbiased look at dataset bias
Antonio Torralba and Alexei A Efros · 2011
Earlier work this paper cites.
The ladder: A reliable leaderboard for machine learning competitions
Avrim Blum and Moritz Hardt · 2015
Earlier work this paper cites.
Binary optimization via mathematical programming with equilibrium constraints
Ganzhao Yuan and Bernard Ghanem · 2016
Earlier work this paper cites.
Cinic-10 is not imagenet or cifar-10
Luke N Darlow, Elliot J Crowley, Antreas Antoniou, and Amos J Storkey · 2018
Earlier work this paper cites.
Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness
Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A Wichmann, and Wieland Brendel · 2018
Earlier work this paper cites.
Do cifar-10 classifiers generalize to cifar-10?
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Earlier work this paper cites.
Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models
Andrei Barbu, David Mayo, Julian Alverio, William Luo, Christopher Wang, Dan Gutfreund, Josh Tenenbaum, and Boris Katz · 2019
Earlier work this paper cites.
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas Dietterich · 2019
Earlier work this paper cites.
Exploring multi-objective exercise recommendations in online education systems
Zhenya Huang, Qi Liu, Chengxiang Zhai, Yu Yin, Enhong Chen, Weibo Gao, and Guoping Hu · 2019
Earlier work this paper cites.
Model similarity mitigates test set overuse
Horia Mania, John Miller, Ludwig Schmidt, Moritz Hardt, and Benjamin Recht · 2019
Earlier work this paper cites.
Do imagenet classifiers generalize to imagenet?
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar · 2019
Earlier work this paper cites.
A meta-analysis of overfitting in machine learning
Rebecca Roelofs, Vaishaal Shankar, Benjamin Recht, Sara Fridovich-Keil, Moritz Hardt, John Miller, and Ludwig Schmidt · 2019
Earlier work this paper cites.
Pytorch image models
Ross Wightman · 2019
Earlier work this paper cites.
The visual task adaptation benchmark
Xiaohua Zhai, Joan Puigcerver, Alexander Kolesnikov, Pierre Ruyssen, Carlos Riquelme, Mario Lucic, Josip Djolonga, Andre Susano Pinto, Maxim Neumann, Alexey Dosovitskiy, et al · 2019
Earlier work this paper cites.
Quality meets diversity: A model-agnostic framework for computerized adaptive testing
Haoyang Bi, Haiping Ma, Zhenya Huang, Yu Yin, Qi Liu, Enhong Chen, Yu Su, and Shijin Wang · 2020
Earlier work this paper cites.
Computing the testing error without a testing set
Ciprian A Corneanu, Sergio Escalera, and Aleix M Martinez · 2020
Earlier work this paper cites.
Evaluating models’ local decision boundaries via contrast sets
Matt Gardner, Yoav Artzi, Victoria Basmova, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, et al · 2020
Earlier work this paper cites.
Beyond accuracy: quantifying trial-by-trial behaviour of cnns and humans by measuring error consistency
Robert Geirhos, Kristof Meding, and Felix A Wichmann · 2020
Earlier work this paper cites.
The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, et al · 2020
Earlier work this paper cites.
Harder or different? a closer look at distribution shift in dataset reproduction
Shangyun Lu, Bradley Nott, Aaron Olson, Alberto Todeschini, Hossein Vahabi, Yair Carmon, and Ludwig Schmidt · 2020
Earlier work this paper cites.
Adversarial nli: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela · 2020
Earlier work this paper cites.
Measuring robustness to natural distribution shifts in image classification
Rohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini, Benjamin Recht, and Ludwig Schmidt · 2020
Earlier work this paper cites.
Deep learning through the lens of example difficulty
Robert Baldock, Hartmut Maennel, and Behnam Neyshabur · 2021
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Earlier work this paper cites.
Are we done with imagenet?
Lucas Beyer, Olivier J Hénaff, Alexander Kolesnikov, Xiaohua Zhai, and Aäron van den Oord · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Earlier work this paper cites.
What will it take to fix benchmarking in natural language understanding?
Samuel R Bowman and George E Dahl · 2021
Earlier work this paper cites.
Mostafa Dehghani, Anurag Arnab, Lucas Beyer, Ashish Vaswani, and Yi Tay · 2021
Earlier work this paper cites.
Nats-bench: Benchmarking nas algorithms for architecture topology and size
Xuanyi Dong, Lu Liu, Katarzyna Musial, and Bogdan Gabrys · 2021
Cited alongside, same era.
Bobcat: Bilevel optimization-based computerized adaptive testing
Aritra Ghosh and Andrew Lan · 2021
Cited alongside, same era.
Active bayesian assessment of black-box classifiers
Disi Ji, Robert L Logan, Padhraic Smyth, and Mark Steyvers · 2021
Cited alongside, same era.
Dynabench: Rethinking benchmarking in nlp
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, et al · 2021
Cited alongside, same era.
Active testing: Sample-efficient model evaluation
Jannik Kossen, Sebastian Farquhar, Yarin Gal, and Tom Rainforth · 2021
Cited alongside, same era.
Are we learning yet? a meta review of evaluation failures across machine learning
Thomas Liao, Rohan Taori, Inioluwa Deborah Raji, and Ludwig Schmidt · 2021
Visit-bench: A benchmark for vision-language instruction following inspired by real-world use
Yonatan Bitton, Hritik Bansal, Jack Hessel, Rulin Shao, Wanrong Zhu, Anas Awadalla, Josh Gardner, Rohan Taori, and Ludwig Schimdt · 2023
Later among the works it cites.
Breaking common sense: Whoops! a vision-and-language benchmark of synthetic and compositional images
Nitzan Bitton-Guetta, Yonatan Bitton, Jack Hessel, Ludwig Schmidt, Yuval Elovici, Gabriel Stanovsky, and Roy Schwartz · 2023
Later among the works it cites.
Pug: Photorealistic and semantically controllable synthetic data for representation learning
Florian Bordes, Shashank Shekhar, Mark Ibrahim, Diane Bouchacourt, Pascal Vincent, and Ari S Morcos · 2023
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Multi-objective optimization of item selection in computerized adaptive testing
Dena F Mujtaba and Nihar R Mahapatra · 2021
Cited alongside, same era.
Valse: A task-independent benchmark for vision and language models centered on linguistic phenomena
Letitia Parcalabescu, Michele Cafagna, Lilitta Muradjan, Anette Frank, Iacer Calixto, and Albert Gatt · 2021
Cited alongside, same era.
Dynasent: A dynamic benchmark for sentiment analysis
Christopher Potts, Zhengxuan Wu, Atticus Geiger, and Douwe Kiela · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Ai and the everything in the whole wide world benchmark
Inioluwa Deborah Raji, Emily M Bender, Amandalynne Paullada, Emily Denton, and Alex Hanna · 2021
Cited alongside, same era.
Evaluation examples are not equally informative: How should that change nlp leaderboards?
Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P Lalor, Robin Jia, and Jordan Boyd-Graber · 2021
Cited alongside, same era.
Muxi Chen, Yu Li, and Qiang Xu · 2023
Later among the works it cites.
Does progress on imagenet transfer to real-world datasets?
Alex Fang, Simon Kornblith, and Ludwig Schmidt · 2023
Later among the works it cites.
Balancing test accuracy and security in computerized adaptive testing
Wanyong Feng, Aritra Ghosh, Stephen Sireci, and Andrew S Lan · 2023
Later among the works it cites.
Datacomp: In search of the next generation of multimodal datasets
Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, et al · 2023
Later among the works it cites.
Adaptive testing of computer vision models
Irena Gao, Gabriel Ilharco, Scott Lundberg, and Marco Tulio Ribeiro · 2023
Later among the works it cites.
Rankme: Assessing the downstream performance of pretrained self-supervised representations by their rank
Quentin Garrido, Randall Balestriero, Laurent Najman, and Yann Lecun · 2023
Later among the works it cites.
Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality
Cheng-Yu Hsieh, Jieyu Zhang, Zixian Ma, Aniruddha Kembhavi, and Ranjay Krishna · 2023
Later among the works it cites.
Llm-perf leaderboard
Régis Pierrard Ilyas Moutawwakil · 2023
Later among the works it cites.
Bring your own data! self-supervised evaluation for large language models
Neel Jain, Khalid Saifullah, Yuxin Wen, John Kirchenbauer, Manli Shu, Aniruddha Saha, Micah Goldblum, Jonas Geiping, and Tom Goldstein · 2023
Later among the works it cites.
Text encoders are performance bottlenecks in contrastive vision-language models
Amita Kamath, Jack Hessel, and Kai-Wei Chang · 2023
Later among the works it cites.
Deconstructing distributions: A pointwise framework of learning
Gal Kaplun, Nikhil Ghosh, Saurabh Garg, Boaz Barak, and Preetum Nakkiran · 2023
Later among the works it cites.
Holistic evaluation of text-to-image models
Tony Lee, Michihiro Yasunaga, Chenlin Meng, Yifan Mai, Joon Sung Park, Agrim Gupta, Yunzhi Zhang, Deepak Narayanan, Hannah Benita Teufel, Marco Bellagente, et al · 2023
Later among the works it cites.
Efficient benchmarking (of language models)
Yotam Perlitz, Elron Bandel, Ariel Gera, Ofir Arviv, Liat Ein-Dor, Eyal Shnarch, Noam Slonim, Michal Shmueli-Scheuer, and Leshem Choshen · 2023
Later among the works it cites.
Automated classification of model errors on imagenet
Momchil Peychev, Mark Niklas Müller, Marc Fischer, and Martin Vechev · 2023
Later among the works it cites.
Beyond chinchilla-optimal: Accounting for inference in language model scaling laws
Nikhil Sardana and Jonathan Frankle · 2023
Later among the works it cites.
Zhelun Shi, Zhipin Wang, Hongxing Fan, Zhenfei Yin, Lu Sheng, Yu Qiao, and Jing Shao · 2023
Later among the works it cites.
What makes imagenet look unlike laion
Ali Shirali and Moritz Hardt · 2023
Later among the works it cites.
Cifar-10-warehouse: Broad and more realistic testbeds in model generalization analysis
Xiaoxiao Sun, Xingjian Leng, Zijian Wang, Yang Yang, Zi Huang, and Liang Zheng · 2023
Later among the works it cites.
Learning vision from models rivals learning vision from data
Yonglong Tian, Lijie Fan, Kaifeng Chen, Dina Katabi, Dilip Krishnan, and Phillip Isola · 2023
Later among the works it cites.
Visual data-type understanding does not emerge from scaling vision-language models
Vishaal Udandarao, Max F Burg, Samuel Albanie, and Matthias Bethge · 2023
Later among the works it cites.
Convnet vs transformer, supervised vs clip: Beyond imagenet accuracy
Kirill Vishniakov, Zhiqiang Shen, and Zhuang Liu · 2023
Later among the works it cites.
Anchor points: Benchmarking models with much fewer examples
Rajan Vivek, Kawin Ethayarajh, Diyi Yang, and Douwe Kiela · 2023
Later among the works it cites.
Gmocat: A graph-enhanced multi-objective method for computerized adaptive testing
Hangyu Wang, Ting Long, Liang Yin, Weinan Zhang, Wei Xia, Qichen Hong, Dingyin Xia, Ruiming Tang, and Yong Yu · 2023
Later among the works it cites.
Sacat: Student-adaptive computerized adaptive testing
Jingwei Yu, Mu Zhenyu, Jiayi Lei, Li’Ang Yin, Wei Xia, Yong Yu, and Ting Long · 2023
Later among the works it cites.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al · 2023
Later among the works it cites.
Model spider: Learning to rank pre-trained models efficiently
Yi-Kai Zhang, Ting-Ji Huang, Yao-Xiang Ding, De-Chuan Zhan, and Han-Jia Ye · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Later among the works it cites.
Lovm: Language-only vision model selection
Orr Zohar, Shih-Cheng Huang, Kuan-Chieh Wang, and Serena Yeung · 2023
Later among the works it cites.
Democratizing evaluation with infinity-benchmarks: Sample-level heterogeneous testing over arbitrary capabilities
Anonymous · 2024
Closest in time.
tinybenchmarks: evaluating llms with fewer examples
Felipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, and Mikhail Yurochkin · 2024
Closest in time.
Flasheval: Towards fast and accurate evaluation of text-to-image diffusion generative models
Lin Zhao, Tianchen Zhao, Zinan Lin, Xuefei Ning, Guohao Dai, Huazhong Yang, and Yu Wang · 2024
Closest in time.