Fetching the paper…
Reading the bibliography…
Attention mechanism is a central component of the transformer architecture which led to the phenomenal success of large language models.
On convergence proofs for perceptrons
Albert B Novikoff · 1963
Earlier work this paper cites.
For valid generalization the size of the weights is more important than the size of the network
Peter Bartlett · 1996
Earlier work this paper cites.
Boosting the margin: A new explanation for the effectiveness of voting methods
Peter Bartlett, Yoav Freund, Wee Sun Lee, and Robert E Schapire · 1998
Earlier work this paper cites.
Margin maximizing loss functions
Saharon Rosset, Ji Zhu, and Trevor Hastie · 2003
Earlier work this paper cites.
Boosting with early stopping: Convergence and consistency
Tong Zhang and Bin Yu · 2005
Earlier work this paper cites.
Estimation of dependences based on empirical data
Vladimir Vapnik · 2006
Earlier work this paper cites.
Margins, shrinkage, and boosting
Matus Telgarsky · 2013
Earlier work this paper cites.
The cifar-10 dataset
Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Long short-term memory-networks for machine reading
Jianpeng Cheng, Li Dong, and Mirella Lapata · 2016
Earlier work this paper cites.
A decomposable attention model for natural language inference
Ankur Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
A structured self-attentive sentence embedding
Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio · 2017
Earlier work this paper cites.
Connecting optimization and regularization paths
Arun Suggala, Adarsh Prasad, and Pradeep K Ravikumar · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Earlier work this paper cites.
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Earlier work this paper cites.
Risk and parameter convergence of logistic regression
Ziwei Ji and Matus Telgarsky · 2018
Earlier work this paper cites.
A deep reinforced model for abstractive summarization
Romain Paulus, Caiming Xiong, and Richard Socher · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Gradient descent aligns the layers of deep linear networks
Ziwei Ji and Matus Telgarsky · 2018
Earlier work this paper cites.
Deep interest network for click-through rate prediction
Guorui Zhou, Xiaoqiang Zhu, Chenru Song, Ying Fan, Han Zhu, Xiao Ma, Yanghui Yan, Junqi Jin, Han Li, and Kun Gai · 2018
Earlier work this paper cites.
The best of both worlds: Combining recent advances in neural machine translation
Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George Foster, Llion Jones, Mike Schuster, Noam Shazeer, Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Zhifeng Chen, Yonghui Wu, and Macduff Hughes · 2018
Earlier work this paper cites.
Universal transformers
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Convergence of gradient descent on separable data
Mor Shpigel Nacson, Jason Lee, Suriya Gunasekar, Pedro Henrique Pamplona Savarese, Nathan Srebro, and Daniel Soudry · 2019
Earlier work this paper cites.
The implicit bias of gradient descent on nonseparable data
Ziwei Ji and Matus Telgarsky · 2019
Earlier work this paper cites.
Implicit regularization for optimal sparse recovery
Tomas Vaskevicius, Varun Kanade, and Patrick Rebeschini · 2019
Earlier work this paper cites.
Yuanzhi Li, Colin Wei, and Tengyu Ma · 2019
Earlier work this paper cites.
The implicit bias of adagrad on separable data
Qian Qian and Xiaoyuan Qian · 2019
Earlier work this paper cites.
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
Samet Oymak and Mahdi Soltanolkotabi · 2019
Earlier work this paper cites.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Earlier work this paper cites.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2019
Earlier work this paper cites.
Behavior sequence transformer for e-commerce recommendation in alibaba
Qiwei Chen, Huan Zhao, Wei Li, Pipei Huang, and Wenwu Ou · 2019
Cited alongside, same era.
Bert4rec: Sequential recommendation with bidirectional encoder representations from transformer
Fei Sun, Jun Liu, Jian Wu, Changhua Pei, Xiao Lin, Wenwu Ou, and Peng Jiang · 2019
Cited alongside, same era.
Are transformers universal approximators of sequence-to-sequence functions?
Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank Reddi, and Sanjiv Kumar · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, and et al · 2020
Cited alongside, same era.
Gradient descent follows the regularization path for general losses
Ziwei Ji, Miroslav Dudík, Robert E Schapire, and Matus Telgarsky · 2020
Cited alongside, same era.
Stochastic mirror descent on overparameterized nonlinear models
Navid Azizan, Sahin Lale, and Babak Hassibi · 2021
Later among the works it cites.
Reinforcement learning as one big sequence modeling problem
Michael Janner, Qiyang Li, and Sergey Levine · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Later among the works it cites.
All tokens matter: Token labeling for training better vision transformers
Zi-Hang Jiang, Qibin Hou, Li Yuan, Daquan Zhou, Yujun Shi, Xiaojie Jin, Anran Wang, and Jiashi Feng · 2021
Later among the works it cites.
The lipschitz constant of self-attention
Hyunjik Kim, George Papamakarios, and Andriy Mnih · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Implicit bias in deep linear classification: Initialization scale vs training accuracy
Edward Moroshko, Blake E Woodworth, Suriya Gunasekar, Jason D Lee, Nati Srebro, and Daniel Soudry · 2020
Cited alongside, same era.
Directional convergence and alignment in deep learning
Ziwei Ji and Matus Telgarsky · 2020
Cited alongside, same era.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Cited alongside, same era.
A unifying view on implicit bias in training linear neural networks
Chulhee Yun, Shankar Krishnan, and Hossein Mobahi · 2020
Cited alongside, same era.
Winnowing with gradient descent
Ehsan Amid and Manfred K Warmuth · 2020
Cited alongside, same era.
Reparameterizing mirror descent as gradient descent
Ehsan Amid and Manfred KK Warmuth · 2020
Cited alongside, same era.
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
Guy Blanc, Neha Gupta, Gregory Valiant, and Paul Valiant · 2020
Cited alongside, same era.
Charlie Snell, Ruiqi Zhong, Dan Klein, and Jacob Steinhardt · 2021
Later among the works it cites.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le · 2021
Later among the works it cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Later among the works it cites.
Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, et al · 2022
Later among the works it cites.
What happens after SGD reaches zero loss? –a mathematical framework
Zhiyuan Li, Tianhao Wang, and Sanjeev Arora · 2022
Later among the works it cites.
Inductive biases and variable creation in self-attention mechanisms
Benjamin L Edelman, Surbhi Goel, Sham Kakade, and Cyril Zhang · 2022
Later among the works it cites.
Unraveling attention via convex duality: Analysis and interpretations of vision transformers
Arda Sahiner, Tolga Ergen, Batu Ozturkler, John Pauly, Morteza Mardani, and Mert Pilanci · 2022
Later among the works it cites.
Convexifying transformers: Improving optimization and understanding of transformer networks
Tolga Ergen, Behnam Neyshabur, and Harsh Mehta · 2022
Later among the works it cites.
Pierre Baldi and Roman Vershynin · 2022
Later among the works it cites.
Vision transformers provably learn spatial structure
Samy Jelassi, Michael Eli Sander, and Yuanzhi Li · 2022
Later among the works it cites.
On margin maximization in linear and relu networks
Gal Vardi, Ohad Shamir, and Nati Srebro · 2022
Later among the works it cites.
Online decision transformer
Qinqing Zheng, Amy Zhang, and Aditya Grover · 2022
Later among the works it cites.
What learning algorithm is in-context learning? investigations with linear models
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2022
Later among the works it cites.
What can transformers learn in-context? a case study of simple function classes
Shivam Garg, Dimitris Tsipras, Percy S Liang, and Gregory Valiant · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou · 2022
Later among the works it cites.
OpenAI · 2023
Closest in time.
Palm-e: An embodied multimodal language model
Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al · 2023
Closest in time.
On the role of attention in prompt-tuning
Samet Oymak, Ankit Singh Rawat, Mahdi Soltanolkotabi, and Christos Thrampoulidis · 2023
Closest in time.
Hongkang Li, Meng Wang, Sijia Liu, and Pin-Yu Chen · 2023
Closest in time.
Mahdi Soltanolkotabi, Dominik Stöger, and Changzhi Xie · 2023
Closest in time.
On generalization of decentralized learning with separable data
Hossein Taheri and Christos Thrampoulidis · 2023
Closest in time.
Benign overfitting in linear classifiers and leaky relu networks from kkt conditions for margin maximization
Spencer Frei, Gal Vardi, Peter L Bartlett, and Nathan Srebro · 2023
Closest in time.
Looped transformers as programmable computers
Angeliki Giannou, Shashank Rajput, Jy-yong Sohn, Kangwook Lee, Jason D Lee, and Dimitris Papailiopoulos · 2023
Closest in time.
Transformers as algorithms: Generalization and stability in in-context learning
Yingcong Li, M Emrullah Ildiz, Dimitris Papailiopoulos, and Samet Oymak · 2023
Closest in time.
Towards revealing the mystery behind chain of thought: a theoretical perspective
Guhao Feng, Yuntian Gu, Bohang Zhang, Haotian Ye, Di He, and Liwei Wang · 2023
Closest in time.
Dissecting chain-of-thought: A study on compositional in-context learning of mlps
Yingcong Li, Kartik Sreenivasan, Angeliki Giannou, Dimitris Papailiopoulos, and Samet Oymak · 2023
Closest in time.
Scan and snap: Understanding training dynamics and token composition in 1-layer transformer
Yuandong Tian, Yiping Wang, Beidi Chen, and Simon Du · 2023
Closest in time.
A primal-dual framework for transformers and neural networks
Tan Minh Nguyen, Tam Minh Nguyen, Nhat Ho, Andrea L Bertozzi, Richard Baraniuk, and Stanley Osher · 2023
Closest in time.
Transformers as support vector machines
Davoud Ataee Tarzanagh, Yingcong Li, Christos Thrampoulidis, and Samet Oymak · 2023
Closest in time.