Fetching the paper…
Reading the bibliography…
The training and generalization dynamics of the Transformer's core mechanism, namely the Attention mechanism, remain under-explored.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Margins, shrinkage, and boosting
Matus Telgarsky · 2013
Earlier work this paper cites.
Long short-term memory-networks for machine reading
Jianpeng Cheng, Li Dong, and Mirella Lapata · 2016
Earlier work this paper cites.
A decomposable attention model for natural language inference
Ankur Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit · 2016
Earlier work this paper cites.
A structured self-attentive sentence embedding
Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Risk and parameter convergence of logistic regression
Ziwei Ji and Matus Telgarsky · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Earlier work this paper cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Yuan Cao and Quanquan Gu · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach, 2019
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Decoupled weight decay regularization, 2019
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Atsushi Nitanda, Geoffrey Chinot, and Taiji Suzuki · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
How much over-parameterization is sufficient to learn deep relu networks?
Zixiang Chen, Yuan Cao, Difan Zou, and Quanquan Gu · 2020
Earlier work this paper cites.
Infinite attention: Nngp and ntk for deep attention networks
Jiri Hron, Yasaman Bahri, Jascha Sohl-Dickstein, and Roman Novak · 2020
Earlier work this paper cites.
Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow relu networks
Ziwei Ji and Matus Telgarsky · 2020
Earlier work this paper cites.
Fine-grained analysis of stability and generalization for stochastic gradient descent
Yunwen Lei and Yiming Ying · 2020
Cited alongside, same era.
On the linearity of large non-linear models: when and why the tangent kernel is constant
Chaoyue Liu, Libin Zhu, and Misha Belkin · 2020
Cited alongside, same era.
Global convergence of deep networks with one wide layer followed by pyramidal topology
Quynh N Nguyen and Marco Mondelli · 2020
Cited alongside, same era.
Toward moderate overparameterization: Global convergence guarantees for training shallow neural networks
Samet Oymak and Mahdi Soltanolkotabi · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Cited alongside, same era.
Convexifying transformers: Improving optimization and understanding of transformer networks
Tolga Ergen, Behnam Neyshabur, and Harsh Mehta · 2022
Later among the works it cites.
Vision transformers provably learn spatial structure
Samy Jelassi, Michael Eli Sander, and Yuanzhi Li · 2022
Later among the works it cites.
Stability and generalization analysis of gradient methods for shallow neural networks
Yunwen Lei, Rong Jin, and Yiming Ying · 2022
Later among the works it cites.
Beyond Lipschitz: Sharp generalization and excess risk bounds for full-batch gd
Konstantinos E Nikolakakis, Farzin Haddadpour, Amin Karbasi, and Dionysios S Kalogerias · 2022
Later among the works it cites.
Openai: Introducing chatgpt, 2022
OpenAI · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale, 2021
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
Inductive biases and variable creation in self-attention mechanisms
Benjamin L Edelman, Surbhi Goel, Sham Kakade, and Cyril Zhang · 2021
Cited alongside, same era.
Characterizing the implicit bias via a primal-dual analysis
Ziwei Ji and Matus Telgarsky · 2021
Cited alongside, same era.
On the expressive power of self-attention matrices, 2021
Valerii Likhosherstov, Krzysztof Choromanski, and Adrian Weller · 2021
Cited alongside, same era.
On the dynamics of training attention models
Haoye Lu, Yongyi Mao, and Amiya Nayak · 2021
Cited alongside, same era.
Tight bounds on the smallest eigenvalue of the neural tangent kernel for deep relu networks
Quynh Nguyen, Marco Mondelli, and Guido F Montufar · 2021
Cited alongside, same era.
Unraveling attention via convex duality: Analysis and interpretations of vision transformers
Arda Sahiner, Tolga Ergen, Batu Ozturkler, John Pauly, Morteza Mardani, and Mert Pilanci · 2022
Later among the works it cites.
Stability vs implicit bias of gradient methods on separable data and beyond
Matan Schliserman and Tomer Koren · 2022
Later among the works it cites.
Feature selection and low test error in shallow low-rotation relu networks
Matus Telgarsky · 2022
Later among the works it cites.
Transformers learn in-context by gradient descent
Johannes von Oswald, Eyvind Niklasson, E. Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2022
Later among the works it cites.
What learning algorithm is in-context learning? investigations with linear models
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2023
Closest in time.
Birth of a transformer: A memory viewpoint
Alberto Bietti, Vivien Cabannes, Diane Bouchacourt, Herve Jegou, and Leon Bottou · 2023
Closest in time.
Memorization capacity of multi-head attention in transformers
Sadegh Mahdavi, Renjie Liao, and Christos Thrampoulidis · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
On the role of attention in prompt-tuning
Samet Oymak, Ankit Singh Rawat, Mahdi Soltanolkotabi, and Christos Thrampoulidis · 2023
Closest in time.
Representational strengths and limitations of transformers
Clayton Sanford, Daniel Hsu, and Matus Telgarsky · 2023
Closest in time.
Generalization and stability of interpolating neural networks with minimal width
Hossein Taheri and Christos Thrampoulidis · 2023
Closest in time.
Scan and snap: Understanding training dynamics and token composition in 1-layer transformer, 2023
Yuandong Tian, Yiping Wang, Beidi Chen, and Simon Du · 2023
Closest in time.
Llama: Open and efficient foundation language models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Closest in time.
Over-parameterization exponentially slows down gradient descent for learning a single neuron
Weihang Xu and Simon Du · 2023
Closest in time.
Trained transformers learn linear models in-context, 2023
Ruiqi Zhang, Spencer Frei, and Peter L. Bartlett · 2023
Closest in time.
Benign overfitting in deep neural networks under lazy training
Zhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Francesco Locatello, and Volkan Cevher · 2023
Closest in time.