On the computational power of transformers and its implications in sequence modeling
Original
Satwik Bhattamishra, Arkil Patel, and Navin Goyal · 2006
Earlier work this paper cites.
On the ability and limitations of transformers to recognize formal languages
Original
Satwik Bhattamishra, Kabir Ahuja, and Navin Goyal · 2009
Earlier work this paper cites.
Pointer sentinel mixture models
Original
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Globally optimal gradient descent for a convnet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Earlier work this paper cites.
When is a convolutional filter easy to learn?
Original
Simon S Du, Jason D Lee, and Yuandong Tian · 2017
Earlier work this paper cites.
Learning relus via gradient descent
Mahdi Soltanolkotabi · 2017
Earlier work this paper cites.
An analytical formula of population gradient for two-layered relu network and its applications in convergence and critical point analysis
Yuandong Tian · 2017
Earlier work this paper cites.
Attention is all you need
Original
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
A convergence analysis of gradient descent for deep linear neural networks
Original
Sanjeev Arora, Nadav Cohen, Noah Golowich, and Wei Hu · 2018
Earlier work this paper cites.
Gradient descent with identity initialization efficiently learns positive definite linear transformations by deep residual networks
Peter Bartlett, Dave Helmbold, and Philip Long · 2018
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lenaic Chizat and Francis Bach · 2018
Earlier work this paper cites.
Universal transformers
Original
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Original
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Learning one convolutional layer with overlapping patches
Surbhi Goel, Adam Klivans, and Raghu Meka · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Earlier work this paper cites.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Earlier work this paper cites.
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Earlier work this paper cites.