Fetching the paper…
Reading the bibliography…
Foundation models are transforming machine learning across many modalities, with in-context learning replacing classical model training.
Learning to control fast-weight memories: An alternative to dynamic recurrent networks
Jürgen Schmidhuber · 1992
Earlier work this paper cites.
Learning to learn: Introduction and overview
Sebastian Thrun and Lorien Pratt · 1998
Earlier work this paper cites.
Random forests
Leo Breiman · 2001
Earlier work this paper cites.
Greedy function approximation: a gradient boosting machine
Jerome H Friedman · 2001
Earlier work this paper cites.
Statistical comparisons of classifiers over multiple data sets
Janez Demšar · 2006
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
James Bergstra, Rémi Bardenet, Yoshua Bengio, and Balázs Kégl · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in python
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al · 2011
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando De Freitas · 2016
Earlier work this paper cites.
Learning feed-forward one-shot learners
Luca Bertinetto, João F Henriques, Jack Valmadre, Philip Torr, and Andrea Vedaldi · 2016
Earlier work this paper cites.
Xgboost: A scalable tree boosting system
Tianqi Chen and Carlos Guestrin · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Optimization as a model for few-shot learning
Sachin Ravi and Hugo Larochelle · 2016
Cited alongside, same era.
Bernd Bischl, Giuseppe Casalicchio, Matthias Feurer, Pieter Gijsbers, Frank Hutter, Michel Lang, Rafael G Mantovani, Jan N van Rijn, and Joaquin Vanschoren · 2017
Cited alongside, same era.
Hypernetworks
David Ha, Andrew Dai, and Quoc V. Le · 2017
Cited alongside, same era.
Lightgbm: A highly efficient gradient boosting decision tree
Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Notes from the ai frontier: Insights from hundreds of use cases
Michael Chui, James Manyika, Mehdi Miremadi, Nicolaus Henke, Rita Chung, Pieter Nel, and Sankalp Malhotra · 2018
Transformers can do bayesian inference
Samuel Müller, Noah Hollmann, Sebastian Pineda Arango, Josif Grabocka, and Frank Hutter · 2021
Later among the works it cites.
Lift: Language-interfaced fine-tuning for non-language machine learning tasks
Tuan Dinh, Yuchen Zeng, Ruisu Zhang, Ziqian Lin, Michael Gira, Shashank Rajput, Jy-yong Sohn, Dimitris Papailiopoulos, and Kangwook Lee · 2022
Later among the works it cites.
TabPFN: A transformer that solves small tabular classification problems in a second
Noah Hollmann, Samuel Müller, Katharina Eggensperger, and Frank Hutter · 2022
Later among the works it cites.
Transformer neural processes: Uncertainty-aware meta learning via sequence modeling
Tung Nguyen and Aditya Grover · 2022
Later among the works it cites.
Efficient transformers: A survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Conditional neural processes
Marta Garnelo, Dan Rosenbaum, Christopher Maddison, Tiago Ramalho, David Saxton, Murray Shanahan, Yee Whye Teh, Danilo Rezende, and SM Ali Eslami · 2018
Cited alongside, same era.
Fourier features let networks learn high frequency functions in low dimensional domains
Matthew Tancik, Pratul Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan Barron, and Ren Ng · 2020
Cited alongside, same era.
Tabert: Pretraining for joint understanding of textual and tabular data
Pengcheng Yin, Graham Neubig, Wen-tau Yih, and Sebastian Riedel · 2020
Cited alongside, same era.
Revisiting deep learning models for tabular data
Yury Gorishniy, Ivan Rubachev, Valentin Khrulkov, and Artem Babenko · 2021
Cited alongside, same era.
Tabbie: Pretrained representations of tabular data
Hiroshi Iida, Dung Thai, Varun Manjunatha, and Mohit Iyyer · 2021
Cited alongside, same era.
Hypertransformer: Model generation for supervised and semi-supervised few-shot learning
Andrey Zhmoginov, Mark Sandler, and Maksym Vladymyrov · 2022
Later among the works it cites.
Hyperfast: Instant classification for tabular data
David Bonet, Daniel Mas Montserrat, Xavier Giró-i Nieto, and Alexander Ioannidis · 2023
Closest in time.
Monarch mixer: A simple sub-quadratic gemm-based architecture
Daniel Y Fu, Simran Arora, Jessica Grogan, Isys Johnson, Sabri Eyuboglu, Armin W Thomas, Benjamin Spector, Michael Poli, Atri Rudra, and Christopher Ré · 2023
Closest in time.
Tabllm: Few-shot classification of tabular data with large language models
Stefan Hegselmann, Alejandro Buendia, Hunter Lang, Monica Agrawal, Xiaoyi Jiang, and David Sontag · 2023
Closest in time.
Hyena hierarchy: Towards larger convolutional language models
Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher Ré · 2023
Closest in time.
When do neural nets outperform boosted trees on tabular data?
Duncan McElfresh, Sujay Khandagale, Jonathan Valverde, Vishak Prasad C, Ganesh Ramakrishnan, Micah Goldblum, and Colin White · 2024
Closest in time.