Fetching the paper…
Reading the bibliography…
The world of empirical machine learning (ML) strongly relies on benchmarks in order to determine the relative effectiveness of different algorithms and methods.
New loss functions for fast maximum inner product search
Ruiqi Guo, Quan Geng, David Simcha, Felix Chern, Sanjiv Kumar, and Xiang Wu · 1908
Earlier work this paper cites.
Computing machinery and intelligence
Alan M. Turing · 1950
Earlier work this paper cites.
Rank aggregation methods for the web
Cynthia Dwork, Ravi Kumar, Moni Naor, and Dandapani Sivakumar · 2001
Earlier work this paper cites.
Learning methods for generic object recognition with invariance to pose and lighting
Yann LeCun, Fu Jie Huang, and Leon Bottou · 2004
Earlier work this paper cites.
Synthesizer: Rethinking self-attention in transformer models
Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng · 2005
Earlier work this paper cites.
Statistical comparisons of classifiers over multiple data sets
Janez Demšar · 2006
Earlier work this paper cites.
One-shot learning of object categories
Li Fei-Fei, Rob Fergus, and Pietro Perona · 2006
Earlier work this paper cites.
Dataset issues in object recognition
Jean Ponce, Tamara L Berg, Mark Everingham, David A Forsyth, Martial Hebert, Svetlana Lazebnik, Marcin Marszalek, Cordelia Schmid, Bryan C Russell, Antonio Torralba, et al · 2006
Earlier work this paper cites.
Automated flower classification over a large number of classes
Maria-Elena Nilsback and Andrew Zisserman · 2008
Earlier work this paper cites.
Improvements that don’t add up: ad-hoc retrieval results since 1998
Timothy G Armstrong, Alistair Moffat, William Webber, and Justin Zobel · 2009
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Efficient transformers: A survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler · 2009
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo
Jianxiong Xiao, James Hays, Krista A Ehinger, Aude Oliva, and Antonio Torralba · 2010
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Long range arena: A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler · 2011
Earlier work this paper cites.
Unbiased look at dataset bias
Antonio Torralba and Alexei A Efros · 2011
Earlier work this paper cites.
Experimental methods for information retrieval
Donald Metzler and Oren Kurland · 2012
Earlier work this paper cites.
Cats and dogs
Omkar M Parkhi, Andrea Vedaldi, Andrew Zisserman, and CV Jawahar · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Vision meets robotics: The kitti dataset
Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Describing textures in the wild
Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
A neuroevolution approach to general atari game playing
Matthew Hausknecht, Joel Lehman, Risto Miikkulainen, and Peter Stone · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Earlier work this paper cites.
Reports on the 2015 aaai workshop program
Stefano V Albrecht, J Christopher Beck, David L Buckeridge, Adi Botea, Cornelia Caragea, Chi-hung Chi, Theodoros Damoulas, Bistra Dilkina, Eric Eaton, Pooyan Fazli, et al · 2015
Earlier work this paper cites.
The ladder: A reliable leaderboard for machine learning competitions
Avrim Blum and Moritz Hardt · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning · 2015
Earlier work this paper cites.
The reusable holdout: Preserving validity in adaptive data analysis
Cynthia Dwork, Vitaly Feldman, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Aaron Roth · 2015
Earlier work this paper cites.
The movielens datasets: History and context
F. Maxwell Harper and Joseph A. Konstan · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Kaggle diabetic retinopathy detection., 2015
Kaggle and EyePacs · 2015
Earlier work this paper cites.
State of the art control of atari games using shallow reinforcement learning
Yitao Liang, Marlos C Machado, Erik Talvitie, and Michael Bowling · 2015
Earlier work this paper cites.
Classical planning with simulators: Results on the atari video games
Nir Lipovetzky, Miquel Ramirez, and Hector Geffner · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Massively parallel methods for deep reinforcement learning
Arun Nair, Praveen Srinivasan, Sam Blackwell, Cagdas Alcicek, Rory Fearon, Alessandro De Maria, Vedavyas Panneershelvam, Mustafa Suleyman, Charles Beattie, Stig Petersen, et al · 2015
Earlier work this paper cites.
Revisiting optimal rank aggregation: A dynamic programming approach
Shayan A Tabrizi, Javid Dadashkarimi, Mostafa Dehghani, Hassan Nasr Esfahani, and Azadeh Shakery · 2015
Earlier work this paper cites.
Charles Beattie, Joel Z Leibo, Denis Teplyashin, Tom Ward, Marcus Wainwright, Heinrich Küttler, Andrew Lefrancq, Simon Green, Víctor Valdés, Amir Sadik, et al · 2016
Earlier work this paper cites.
Real-world electronic voting: Design, analysis and deployment
Feng Hao and Peter YA Ryan · 2016
Earlier work this paper cites.
Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering
Ruining He and Julian McAuley · 2016
Earlier work this paper cites.
beta-vae: Learning basic visual concepts with a constrained variational framework
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner · 2016
Earlier work this paper cites.
Reinforcement learning with unsupervised auxiliary tasks
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z Leibo, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Assessing the ability of lstms to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
A benchmark dataset and evaluation methodology for video object segmentation
Federico Perazzi, Jordi Pont-Tuset, Brian McWilliams, Luc Van Gool, Markus Gross, and Alexander Sorkine-Hornung · 2016
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Off-policy evaluation for slate recommendation
Adith Swaminathan, Akshay Krishnamurthy, Alekh Agarwal, Miroslav Dudík, John Langford, Damien Jose, and Imed Zitouni · 2016
Do ImageNet classifiers generalize to ImageNet?
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar · 2019
Later among the works it cites.
A meta-analysis of overfitting in machine learning
Rebecca Roelofs, Vaishaal Shankar, Benjamin Recht, Sara Fridovich-Keil, Moritz Hardt, John Miller, and Ludwig Schmidt · 2019
Later among the works it cites.
The evolved transformer
David So, Quoc Le, and Chen Liang · 2019
Later among the works it cites.
Best practices for the human evaluation of automatically generated text
Chris Van Der Lee, Albert Gatt, Emiel Van Miltenburg, Sander Wubben, and Emiel Krahmer · 2019
Later among the works it cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Hasselt, Marc Lanctot, and Nando Freitas · 2016
Cited alongside, same era.
Beyond globally optimal: Focused learning for improved recommendations
Alex Beutel, Ed H Chi, Zhiyuan Cheng, Hubert Pham, and John Anderson · 2017
Cited alongside, same era.
Remote sensing image scene classification: Benchmark and state of the art
Gong Cheng, Junwei Han, and Xiaoqiang Lu · 2017
Cited alongside, same era.
Replicability analysis for natural language processing: Testing significance with multiple datasets
Rotem Dror, Gili Baumer, Marina Bogomolov, and Roi Reichart · 2017
Cited alongside, same era.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith · 2017
Cited alongside, same era.
Neural collaborative filtering
Xiangnan He, Lizi Liao, Hanwang Zhang, Liqiang Nie, Xia Hu, and Tat-Seng Chua · 2017
Cited alongside, same era.
Chhavi Yadav and Léon Bottou · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le · 2019
Later among the works it cites.
Sampling-bias-corrected neural modeling for large corpus item recommendations
Xinyang Yi, Ji Yang, Lichan Hong, Derek Zhiyuan Cheng, Lukasz Heldt, Aditee Kumthekar, Zhe Zhao, Li Wei, and Ed Chi · 2019
Later among the works it cites.
A large-scale study of representation learning with the visual task adaptation benchmark
Xiaohua Zhai, Joan Puigcerver, Alexander Kolesnikov, Pierre Ruyssen, Carlos Riquelme, Mario Lucic, Josip Djolonga, Andre Susano Pinto, Maxim Neumann, Alexey Dosovitskiy, et al · 2019
Later among the works it cites.
Deep learning based recommender system: A survey and new perspectives
Shuai Zhang, Lina Yao, Aixin Sun, and Yi Tay · 2019
Later among the works it cites.
Recommending what video to watch next: a multitask ranking system
Zhe Zhao, Lichan Hong, Li Wei, Jilin Chen, Aniruddh Nath, Shawn Andrews, Aditee Kumthekar, Maheswaran Sathiamoorthy, Xinyang Yi, and Ed Chi · 2019
Later among the works it cites.
Mars: Memory attention-aware recommender system
Lei Zheng, Chun-Ta Lu, Lifang He, Sihong Xie, Huang He, Chaozhuo Li, Vahid Noroozi, Bowen Dong, and S Yu Philip · 2019
Later among the works it cites.
Transferring inductive biases through knowledge distillation
Samira Abnar, Mostafa Dehghani, and Willem Zuidema · 2020
Later among the works it cites.
Identifying annotator bias: A new irt-based method for bias identification
Jacopo Amidei, Paul Piwek, and Alistair Willis · 2020
Later among the works it cites.
Lucas Beyer, Olivier J Hénaff, Alexander Kolesnikov, Xiaohua Zhai, and Aäron van den Oord · 2020
Later among the works it cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Electra: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning · 2020
Later among the works it cites.
Bringing the people back in: Contesting benchmark machine learning datasets
Emily Denton, Alex Hanna, Razvan Amironesei, Andrew Smart, Hilary Nicole, and Morgan Klaus Scheuerman · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Later among the works it cites.
Utility is in the eye of the user: A critique of nlp leaderboards
Kawin Ethayarajh and Dan Jurafsky · 2020
Later among the works it cites.
Rl unplugged: A collection of benchmarks for offline reinforcement learning
Caglar Gulcehre, Ziyu Wang, Alexander Novikov, Thomas Paine, Sergio Gómez, Konrad Zolna, Rishabh Agarwal, Josh S Merel, Daniel J Mankowitz, Cosmin Paduraru, et al · 2020
Later among the works it cites.
The many faces of robustness: A critical analysis of out-of-distribution generalization
Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt, and Justin Gilmer · 2020
Later among the works it cites.
Sara Hooker · 2020
Later among the works it cites.
Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson · 2020
Later among the works it cites.
A metric learning reality check
Kevin Musgrave, Serge Belongie, and Ser-Nam Lim · 2020
Later among the works it cites.
Neural collaborative filtering vs. matrix factorization revisited
Steffen Rendle, Walid Krichene, Li Zhang, and John Anderson · 2020
Later among the works it cites.
“this is a problem, don’t you agree?” framing and bias in human evaluation for natural language generation
Stephanie Schoch, Diyi Yang, and Yangfeng Ji · 2020
Later among the works it cites.
Self-supervised learning of video-induced visual invariances
Michael Tschannen, Josip Djolonga, Marvin Ritter, Aravindh Mahendran, Neil Houlsby, Sylvain Gelly, and Mario Lucic · 2020
Later among the works it cites.
Comparing test sets with item response theory
Clara Vania, Phu Mon Htut, William Huang, Dhara Mungra, Richard Yuanzhe Pang, Jason Phang, Haokun Liu, Kyunghyun Cho, and Samuel R. Bowman · 2020
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding, 2020
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V. Le · 2020
Later among the works it cites.
A model of two tales: Dual transfer learning framework for improved long-tail item recommendation
Yin Zhang, Derek Zhiyuan Cheng, Tiansheng Yao, Xinyang Yi, Lichan Hong, and Ed H Chi · 2020
Later among the works it cites.
Rip van winkle’s razor: A simple estimate of overfit to test data
Sanjeev Arora and Yi Zhang · 2021
Closest in time.
Revisiting resnets: Improved training and scaling strategies
Irwan Bello, William Fedus, Xianzhi Du, Ekin D Cubuk, Aravind Srinivas, Tsung-Yi Lin, Jonathon Shlens, and Barret Zoph · 2021
Closest in time.
Accounting for variance in machine learning benchmarks
Xavier Bouthillier, Pierre Delaunay, Mirko Bronzi, Assya Trofimov, Brennan Nichyporuk, Justin Szeto, Nazanin Mohammadi Sepahvand, Edward Raff, Kanika Madan, Vikram Voleti, et al · 2021
Closest in time.
What will it take to fix benchmarking in natural language understanding?
Samuel R Bowman and George E Dahl · 2021
Closest in time.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Closest in time.
The GEM benchmark: Natural language generation, its evaluation and metrics
Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Aremu Anuoluwapo, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna Clinciu, Dipanjan Das, Kaustubh D Dhole, et al · 2021
Closest in time.
Natural adversarial examples
Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song · 2021
Closest in time.
Dynabench: Rethinking benchmarking in nlp
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, et al · 2021
Closest in time.
Jimmy Lin, Daniel Campos, Nick Craswell, Bhaskar Mitra, and Emine Yilmaz · 2021
Closest in time.
Scientific credibility of machine translation research: A meta-evaluation of 769 papers
Benjamin Marie, Atsushi Fujita, and Raphael Rubino · 2021
Closest in time.
How robust are model rankings: A leaderboard customization approach for equitable evaluation
Swaroop Mishra and Anjana Arunkumar · 2021
Closest in time.
Do transformer modifications transfer across implementations and applications?, 2021
Sharan Narang, Hyung Won Chung, Yi Tay, William Fedus, Thibault Fevry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam Shazeer, Zhenzhong Lan, Yanqi Zhou, Wei Li, Nan Ding, Jake Marcus, Adam Roberts, and Colin Raffel · 2021
Closest in time.
Pervasive label errors in test sets destabilize machine learning benchmarks
Curtis G Northcutt, Anish Athalye, and Jonas Mueller · 2021
Closest in time.
Online and offline reinforcement learning by planning with a learned model
Julian Schrittwieser, Thomas Hubert, Amol Mandhane, Mohammadamin Barekatain, Ioannis Antonoglou, and David Silver · 2021
Closest in time.
How to train your vit? data, augmentation, and regularization in vision transformers
Andreas Steiner, Alexander Kolesnikov, Xiaohua Zhai, Ross Wightman, Jakob Uszkoreit, and Lucas Beyer · 2021
Closest in time.
Are pre-trained convolutions better than pre-trained transformers?
Yi Tay, Mostafa Dehghani, Jai Gupta, Dara Bahri, Vamsi Aribandi, Zhen Qin, and Donald Metzler · 2021
Closest in time.
The need for performance based assessments, 2021
Tianhao Zhang · 2021
Closest in time.