Fetching the paper…
Reading the bibliography…
We challenge a common assumption underlying most supervised deep learning: that a model makes a prediction depending only on its parameters and the features of a single input.
Problems in the analysis of survey data, and a proposal
James N Morgan and John A Sonquist · 1963
Earlier work this paper cites.
Multidimensional binary search trees used for associative searching
J. L. Bentley · 1975
Earlier work this paper cites.
Classification and regression trees
Leo Breiman, Jerome Friedman, Charles J Stone, and Richard A Olshen · 1984
Earlier work this paper cites.
The role of metalearning in study processes
John B Biggs · 1985
Earlier work this paper cites.
Discriminatory analysis: nonparametric discrimination, consistency properties , volume 1
Evelyn Fix · 1985
Earlier work this paper cites.
The strength of weak learnability
Robert E Schapire · 1990
Earlier work this paper cites.
Learning a synaptic learning rule
Y. Bengio, S. Bengio, and J. Cloutier · 1991
Earlier work this paper cites.
An introduction to kernel and nearest-neighbor nonparametric regression
Naomi S Altman · 1992
Earlier work this paper cites.
Bagging predictors
Leo Breiman · 1996
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E Schapire · 1997
Earlier work this paper cites.
Local dimensionality reduction for locally weighted learning
Sethu Vijayakumar and Stefan Schaal · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Random forests
Leo Breiman · 2001
Earlier work this paper cites.
Greedy function approximation: a gradient boosting machine
Jerome H Friedman · 2001
Earlier work this paper cites.
Analyzing incomplete political science data: An alternative algorithm for multiple imputation
Gary King, James Honaker, Anne Joseph, and Kenneth Scheve · 2001
Earlier work this paper cites.
Gaussian processes in machine learning
Carl Edward Rasmussen · 2003
Earlier work this paper cites.
Neighbourhood component analysis
Sam Roweis, Geoffrey Hinton, and Ruslan Salakhutdinov · 2004
Earlier work this paper cites.
New algorithms for efficient high-dimensional nonparametric classification
T. Liu, A. Moore, and A. Gray · 2006
Earlier work this paper cites.
Estimation of dependences based on empirical data
Vladimir Vapnik · 2006
Earlier work this paper cites.
Semi-supervised learning
Olivier Chapelle, Bernhard Scholkopf, and Alexander Zien · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images, 2009
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Why does unsupervised pre-training help deep learning?
Dumitru Erhan, Aaron Courville, Yoshua Bengio, and Pascal Vincent · 2010
Earlier work this paper cites.
What to do about missing values in time series cross-section data
James Honaker and Gary King · 2010
Earlier work this paper cites.
Mnist handwritten digit database
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al · 2011
Earlier work this paper cites.
mice: Multivariate imputation by chained equations in r
Stef van Buuren and Karin Groothuis-Oudshoorn · 2011
Earlier work this paper cites.
Missforest - nonparametric missing value imputation for mixed-type data
D.J. Stekhoven and P. Buehlmann · 2012
Earlier work this paper cites.
Multiple imputation with diagnostics (mi) in R: Opening windows into the black box
Yu-Sung Su, Andrew E. Gelman, Jennifer Hill, and Masanao Yajima · 2012
Earlier work this paper cites.
Deep gaussian processes
Andreas Damianou and Neil D Lawrence · 2013
Earlier work this paper cites.
Semi-supervised learning with deep generative models
Diederik P Kingma, Danilo J Rezende, Shakir Mohamed, and Max Welling · 2014
Earlier work this paper cites.
Fifty years of classification and regression trees
Wei-Yin Loh · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Cited alongside, same era.
Human-level concept learning through probabilistic program induction
Brenden M. Lake, Ruslan Salakhutdinov, and Joshua B. Tenenbaum · 2015
Cited alongside, same era.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Cited alongside, same era.
Deep gaussian processes for regression using approximate expectation propagation
Thang Bui, Daniel Hernández-Lobato, Jose Hernandez-Lobato, Yingzhen Li, and Richard Turner · 2016
Cited alongside, same era.
Xgboost: A scalable tree boosting system
Tianqi Chen and Carlos Guestrin · 2016
Cited alongside, same era.
Variational auto-encoded deep gaussian processes
Axial attention in multidimensional transformers
Jonathan Ho, Nal Kalchbrenner, Dirk Weissenborn, and Tim Salimans · 2019
Later among the works it cites.
Attentive neural processes
Hyunjik Kim, Andriy Mnih, Jonathan Schwarz, Marta Garnelo, Ali Eslami, Dan Rosenbaum, Oriol Vinyals, and Yee Whye Teh · 2019
Later among the works it cites.
Set transformer: A framework for attention-based permutation-invariant neural networks
Juho Lee, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Seungjin Choi, and Yee Whye Teh · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhenwen Dai, Andreas Damianou, Javier González, and Neil Lawrence · 2016
Cited alongside, same era.
Matching networks for one shot learning
Oriol Vinyals, Charles Blundell, Timothy Lillicrap, Daan Wierstra, et al · 2016
Cited alongside, same era.
Deep kernel learning
Andrew Gordon Wilson, Zhiting Hu, Ruslan Salakhutdinov, and Eric P. Xing · 2016
Cited alongside, same era.
UCI machine learning repository, 2017
Dheeru Dua and Casey Graff · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Lightgbm: A highly efficient gradient boosting decision tree
Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu · 2017
Cited alongside, same era.
Semi-supervised classification with graph convolutional networks
Thomas Kipf and Max Welling · 2017
Cited alongside, same era.
Later among the works it cites.
Autoint: Automatic feature interaction learning via self-attentive neural networks
Weiping Song, Chence Shi, Zhiping Xiao, Zhijian Duan, Yewen Xu, Ming Zhang, and Jian Tang · 2019
Later among the works it cites.
Fixing the train-test resolution discrepancy
Hugo Touvron, Andrea Vedaldi, Matthijs Douze, and Herve Jegou · 2019
Later among the works it cites.
Ranked list loss for deep metric learning
Xinshao Wang, Yang Hua, Elyor Kodirov, Guosheng Hu, Romain Garnier, and Neil M Robertson · 2019
Later among the works it cites.
How powerful are graph neural networks?
Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka · 2019
Later among the works it cites.
Lookahead optimizer: k steps forward, 1 step back
Michael Zhang, James Lucas, Jimmy Ba, and Geoffrey E Hinton · 2019
Later among the works it cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan · 2020
Later among the works it cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2020
Later among the works it cites.
Se(3)-transformers: 3d roto-translation equivariant attention networks
Fabian Fuchs, Daniel Worrall, Volker Fischer, and Max Welling · 2020
Later among the works it cites.
A unified view of label shift estimation
Saurabh Garg, Yifan Wu, Sivaraman Balakrishnan, and Zachary Lipton · 2020
Later among the works it cites.
Realm: Retrieval-augmented language model pre-training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang · 2020
Later among the works it cites.
Array programming with NumPy
Charles R. Harris, K. Jarrod Millman, Stéfan J. van der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J. Smith, Robert Kern, Matti Picus, Stephan Hoyer, Marten H. van Kerkwijk, Matthew Brett, Allan Haldane, Jaime Fernández del Río, Mark Wiebe, Pearu Peterson, Pierre Gérard-Marchant, Kevin Sheppard, Tyler Reddy, Warren Weckesser, Hameer Abbasi, Christoph Gohlke, and Travis E. Oliphant · 2020
Later among the works it cites.
Lietransformer: Equivariant self-attention for lie groups
Michael Hutchinson, Charline Le Lan, Sheheryar Zaidi, Emilien Dupont, Yee Whye Teh, and Hyunjik Kim · 2020
Later among the works it cites.
Transformers are RNNs: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Later among the works it cites.
Big transfer (bit): General visual representation learning
Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby · 2020
Later among the works it cites.
Object-centric learning with slot attention
Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf · 2020
Later among the works it cites.
Efficient transformers: A survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler · 2020
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2020
Later among the works it cites.
Large batch optimization for deep learning: Training bert in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2020
Later among the works it cites.
Rethinking attention with performers
Krzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Quincy Davis, Afroz Mohiuddin, Lukasz Kaiser, David Benjamin Belanger, Lucy J Colwell, and Adrian Weller · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Closest in time.
Perceiver: General perception with iterative attention
Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and Joao Carreira · 2021
Closest in time.
Do transformer modifications transfer across implementations and applications?
Sharan Narang, Hyung Won Chung, Yi Tay, William Fedus, Thibault Févry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam Shazeer, Zhenzhong Lan, Yanqi Zhou, Wei Li, Nan Ding, Jake Marcus, Adam Roberts, and Colin Raffel · 2021
Closest in time.
Getting started with the built-in tabnet algorithm, 2021
Google Cloud AI Platform · 2021
Closest in time.
Msa transformer
Roshan Rao, Jason Liu, Robert Verkuil, Joshua Meier, John F Canny, Pieter Abbeel, Tom Sercu, and Alexander Rives · 2021
Closest in time.
Imagenet-21k pretraining for the masses
Tal Ridnik, Emanuel Ben-Baruch, Asaf Noy, and Lihi Zelnik-Manor · 2021
Closest in time.
Learning intra-batch connections for deep metric learning
Jenny Seidenschwarz, Ismail Elezi, and Laura Leal-Taixé · 2021
Closest in time.