Fetching the paper…
Reading the bibliography…
The neural tangent kernel (NTK) has garnered significant attention as a theoretical framework for describing the behavior of large-scale neural networks.
Truth or Backpropaganda? An Empirical Investigation of Deep Learning Theory
Micah Goldblum, Jonas Geiping, Avi Schwarzschild, Michael Moeller, and Tom Goldstein · 1910
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen · 1989
Earlier work this paper cites.
The evidence framework applied to classification networks
David JC MacKay · 1992
Earlier work this paper cites.
Lifelong robot learning
Sebastian Thrun and Tom M Mitchell · 1995
Earlier work this paper cites.
ARPACK users’ guide: solution of large-scale eigenvalue problems with implicitly restarted Arnoldi methods
Richard B Lehoucq, Danny C Sorensen, and Chao Yang · 1998
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2002
Earlier work this paper cites.
Gaussian processes in machine learning
Carl Edward Rasmussen and Christopher K. I. Williams · 2005
Earlier work this paper cites.
Generalisation guarantees for continual learning with orthogonal gradient descent
Mehdi Abbana Bennani, Thang Doan, and Masashi Sugiyama · 2006
Earlier work this paper cites.
Numerical optimization
Jorge Nocedal and Stephen Wright · 2006
Earlier work this paper cites.
Finite Versus Infinite Neural Networks: an Empirical Study
Jaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam, Lechao Xiao, Roman Novak, and Jascha Sohl-Dickstein · 2007
Earlier work this paper cites.
Accelerating the cubic regularization of newton’s method on convex problems
Yu Nesterov · 2008
Earlier work this paper cites.
Feature learning in infinite-width neural networks
Greg Yang and Edward J. Hu · 2011
Earlier work this paper cites.
Contextual Gaussian process bandit optimization
Andreas Krause and Cheng Ong · 2011
Earlier work this paper cites.
An empirical investigation of catastrophic forgetting in gradient-based neural networks
Ian J Goodfellow, Mehdi Mirza, Da Xiao, Aaron Courville, and Yoshua Bengio · 2013
Earlier work this paper cites.
Optimizing neural networks with Kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Wide residual networks (WideResNets) in PyTorch
Jason Kuen · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Deep neural networks as Gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Earlier work this paper cites.
A Mean Field View of the Landscape of Two-Layers Neural Networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Utilizing BERT for aspect-based sentiment analysis via constructing auxiliary sentence
Chi Sun, Luyao Huang, and Xipeng Qiu · 2019
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Precise characterization of the prior predictive distribution of deep ReLU networks
Lorenzo Noci, Gregor Bachmann, Kevin Roth, Sebastian Nowozin, and Thomas Hofmann · 2021
Later among the works it cites.
Tensor programs IIb: Architectural universality of neural tangent kernel training dynamics
Greg Yang and Etai Littwin · 2021
Later among the works it cites.
Deep networks and the multiple manifold problem
Sam Buchanan, Dar Gilboa, and John Wright · 2021
Later among the works it cites.
Superfast second-order methods for unconstrained convex optimization
Yurii Nesterov · 2021
Later among the works it cites.
Neural Thompson sampling
Weitong Zhang, Dongruo Zhou, Lihong Li, and Quanquan Gu · 2021
Later among the works it cites.
Laplace redux–effortless Bayesian deep learning
Erik Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen, Matthias Bauer, and Philipp Hennig · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Russ R Salakhutdinov, and Ruosong Wang · 2019
Cited alongside, same era.
Why ReLU networks yield high-confidence predictions far away from the training data and how to mitigate the problem
Matthias Hein, Maksym Andriushchenko, and Julian Bitterwolf · 2019
Cited alongside, same era.
Approximate inference turns deep networks into Gaussian processes
Mohammad Emtiyaz E Khan, Alexander Immer, Ehsan Abedi, and Maciej Korzepa · 2019
Cited alongside, same era.
Fast convergence of natural gradient descent for over-parameterized neural networks
Guodong Zhang, James Martens, and Roger B Grosse · 2019
Cited alongside, same era.
Continual lifelong learning with neural networks: A review
German I Parisi, Ronald Kemker, Jose L Part, Christopher Kanan, and Stefan Wermter · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Neural contextual bandits with UCB-based exploration
Dongruo Zhou, Lihong Li, and Quanquan Gu · 2020
Cited alongside, same era.
Later among the works it cites.
Avalanche: an end-to-end library for continual learning
Vincenzo Lomonaco, Lorenzo Pellegrini, Andrea Cossu, Antonio Carta, Gabriele Graffieti, Tyler L. Hayes, Matthias De Lange, Marc Masana, Jary Pomponi, Gido van de Ven, Martin Mundt, Qi She, Keiland Cooper, Jeremy Forest, Eden Belouadah, Simone Calderara, German I. Parisi, Fabio Cuzzolin, Andreas Tolias, Simone Scardapane, Luca Antiga, Subutai Amhad, Adrian Popescu, Christopher Kanan, Joost van de Weijer, Tinne Tuytelaars, Davide Bacciu, and Davide Maltoni · 2021
Later among the works it cites.
Wide neural networks forget less catastrophically
Seyed Iman Mirzadeh, Arslan Chaudhry, Dong Yin, Huiyi Hu, Razvan Pascanu, Dilan Gorur, and Mehrdad Farajtabar · 2022
Later among the works it cites.
Effect of scale on catastrophic forgetting in neural networks
Vinay Venkatesh Ramasesh, Aitor Lewkowycz, and Ethan Dyer · 2022
Later among the works it cites.
How catastrophic can catastrophic forgetting be in linear regression?
Itay Evron, Edward Moroshko, Rachel Ward, Nathan Srebro, and Daniel Soudry · 2022
Later among the works it cites.
Wide mean-field Bayesian neural networks ignore the data
Beau Coker, Wessel P. Bruinsma, David R. Burt, Weiwei Pan, and Finale Doshi-Velez · 2022
Later among the works it cites.
Neural contextual bandits without regret
Parnian Kassraie and Andreas Krause · 2022
Later among the works it cites.
Offline neural contextual bandits: Pessimism, optimization and generalization
Thanh Nguyen-Tang, Sunil Gupta, A. Tuan Nguyen, and Svetha Venkatesh · 2022
Later among the works it cites.
Expected improvement for contextual bandits
Sunil Gupta, Santu Rana, Tuan Truong, Long Tran-Thanh, Svetha Venkatesh, et al · 2022
Later among the works it cites.
LLaMA: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
A study of Bayesian neural network surrogates for Bayesian optimization
Yucen Lily Li, Tim GJ Rudner, and Andrew Gordon Wilson · 2023
Closest in time.
Continual learning in linear classification on separable data
Itay Evron, Edward Moroshko, Gon Buzaglo, Maroun Khriesh, Badea Marjieh, Nathan Srebro, and Daniel Soudry · 2023
Closest in time.
Analysis of Catastrophic Forgetting for Random Orthogonal Transformation Tasks in the Overparameterized Regime
Daniel Goldfarb and Paul Hand · 2023
Closest in time.
Empirical Limitations of the NTK for Understanding Scaling Laws in Deep Learning
Nikhil Vyas, Yamini Bansal, and Preetum Nakkiran · 2023
Closest in time.
Bayesian optimization
Roman Garnett · 2023
Closest in time.