Fetching the paper…
Reading the bibliography…
Over the past few years, there has been a significant amount of research focused on studying the ReLU activation function, with the aim of achieving neural network convergence through over-parametrization.
On a modification of chebyshev’s inequality and of the error formula of laplace
Sergei Bernstein · 1924
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Wassily Hoeffding · 1963
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
Beatrice Laurent and Pascal Massart · 2000
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser · 2018
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
On the convergence rate of training recurrent neural networks
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
A gram-gauss-newton method learning overparameterized deep neural networks for regression problems
Tianle Cai, Ruiqi Gao, Jikai Hou, Siyu Chen, Dong Wang, Di He, Zhihua Zhang, and Liwei Wang · 2019
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Probabilistic neural network with complex exponential activation functions in image recognition
Andrey V Savchenko · 2019
Earlier work this paper cites.
Quadratic suffices for over-parametrization via matrix chernoff bound
Zhao Song and Xin Yang · 2019
Earlier work this paper cites.
Fast convergence of natural gradient descent for over-parameterized neural networks
Guodong Zhang, James Martens, and Roger B Grosse · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Kernel density estimation through density constrained near neighbor search
Moses Charikar, Michael Kapralov, Navid Nouri, and Paris Siminelakis · 2020
Cited alongside, same era.
Infinite attention: Nngp and ntk for deep attention networks
Jiri Hron, Yasaman Bahri, Jascha Sohl-Dickstein, and Roman Novak · 2020
Cited alongside, same era.
A mathematical theory of attention
Vuckovic James, Baratin Aristide, and Remi Tachet des Combes · 2020
Cited alongside, same era.
A sublinear adversarial training algorithm
Yeqi Gao, Lianke Qin, Zhao Song, and Yitan Wang · 2022
Later among the works it cites.
What can transformers learn in-context? a case study of simple function classes
Shivam Garg, Dimitris Tsipras, Percy Liang, and Gregory Valiant · 2022
Later among the works it cites.
Training overparametrized neural networks in sublinear time
Hang Hu, Zhao Song, Omri Weinstein, and Danyang Zhuo · 2022
Later among the works it cites.
Transformers learn shortcuts to automata
Bingbin Liu, Jordan T Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang · 2022
Later among the works it cites.
Bounding the width of neural networks via coupled initialization a worst case analysis
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Toward moderate overparameterization: Global convergence guarantees for training shallow neural networks
Samet Oymak and Mahdi Soltanolkotabi · 2020
Cited alongside, same era.
Why are adaptive methods good for attention models?
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank Reddi, Sanjiv Kumar, and Suvrit Sra · 2020
Cited alongside, same era.
Training (overparametrized) neural networks in near-linear time
Jan van den Brand, Binghui Peng, Zhao Song, and Omri Weinstein · 2021
Cited alongside, same era.
Does preprocessing help training over-parameterized neural networks?
Zhao Song, Shuo Yang, and Ruizhe Zhang · 2021
Cited alongside, same era.
Approximating how single head attention learns
Charlie Snell, Ruiqi Zhong, Dan Klein, and Jacob Steinhardt · 2021
Cited alongside, same era.
Training multi-layer over-parametrized neural network in subquadratic time
Zhao Song, Lichen Zhang, and Ruizhe Zhang · 2021
Cited alongside, same era.
An explanation of in-context learning as implicit bayesian inference
Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma · 2021
Cited alongside, same era.
Alexander Munteanu, Simon Omlor, Zhao Song, and David Woodruff · 2022
Later among the works it cites.
Saturated transformers are constant-depth threshold circuits
William Merrill, Ashish Sabharwal, and Noah A Smith · 2022
Later among the works it cites.
Transformers learn in-context by gradient descent
Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2022
Later among the works it cites.
Statistically meaningful approximation: a case study on approximating turing machines with transformers
Colin Wei, Yining Chen, and Tengyu Ma · 2022
Later among the works it cites.
Speeding up optimizations via data structures: Faster search, sample and maintenance
Lichen Zhang · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Later among the works it cites.
Fast attention requires bounded entries
Josh Alman and Zhao Song · 2023
Closest in time.
Algorithms and hardness for dynamic attention maintenance in large language models
Jan van den Brand, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
Looped transformers as programmable computers
Angeliki Giannou, Shashank Rajput, Jy-yong Sohn, Kangwook Lee, Jason D Lee, and Dimitris Papailiopoulos · 2023
Closest in time.
A watermark for large language models
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein · 2023
Closest in time.
How do transformers learn topic structure: Towards a mechanistic understanding
Yuchen Li, Yuanzhi Li, and Andrej Risteski · 2023
Closest in time.
Solving regularized exp, cosh and sinh regression problems
Zhihang Li, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Adaptive representations for semantic search
Aniket Rege, Aditya Kusupati, Sharan Ranjit, Sham Kakade, Prateek Jain, and Ali Farhadi · 2023
Closest in time.
Provable copyright protection for generative models
Nikhil Vyas, Sham Kakade, and Boaz Barak · 2023
Closest in time.
Kdeformer: Accelerating transformers via kernel density estimation
Amir Zandieh, Insu Han, Majid Daliri, and Amin Karbasi · 2023
Closest in time.