Fetching the paper…
Reading the bibliography…
The multi-head attention layer is one of the key components of the transformer architecture that sets it apart from traditional feed-forward models.
Distance matrix polynomials of trees
Ronald L Graham and Laszlo Lovasz · 1978
Earlier work this paper cites.
Concentration of measure and isoperimetric inequalities in product spaces
Michel Talagrand · 1995
Earlier work this paper cites.
Exponential integrability and transportation cost related to logarithmic sobolev inequalities
Sergej G Bobkov and Friedrich Götze · 1999
Earlier work this paper cites.
Lower bounds on large deviation probabilities for sums of independent random variables
Sergey V Nagaev · 2002
Earlier work this paper cites.
Concentration inequalities using the entropy method
Stéphane Boucheron, Gábor Lugosi, and Pascal Massart · 2003
Earlier work this paper cites.
Exponential decay of entropy in the random transposition and bernoulli-laplace models
Fuqing Gao and Jeremy Quastel · 2003
Earlier work this paper cites.
On concentration of self-bounding functions
Stephane Boucheron, Gabor Lugosi, and Pascal Massart · 2009
Earlier work this paper cites.
On lattices, learning with errors, random linear codes, and cryptography
Oded Regev · 2009
Earlier work this paper cites.
Characterizing statistical query learning: simplified notions and proofs
Balázs Szörényi · 2009
Earlier work this paper cites.
Pseudorandom functions and lattices
Abhishek Banerjee, Chris Peikert, and Alon Rosen · 2012
Earlier work this paper cites.
Learning with rounding, revisited: New reduction, properties and applications
Joël Alwen, Stephan Krenn, Krzysztof Pietrzak, and Daniel Wichs · 2013
Earlier work this paper cites.
Concentration inequalities: A nonasymptotic theory of independence
Stéphane Boucheron, Gábor Lugosi, and Pascal Massart · 2013
Earlier work this paper cites.
A note on the hanson-wright inequality for random vectors with dependencies
Radosław Adamczak · 2014
Earlier work this paper cites.
The convex distance inequality for dependent random variables, with applications to the stochastic travelling salesman and other problems
Daniel Paulin · 2014
Earlier work this paper cites.
On the hardness of learning with rounding over small modulus
Andrej Bogdanov, Siyao Guo, Daniel Masny, Silas Richelson, and Alon Rosen · 2015
Earlier work this paper cites.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
Majid Janzamin, Hanie Sedghi, and Anima Anandkumar · 2015
Earlier work this paper cites.
Provable tensor methods for learning mixtures of generalized linear models
Hanie Sedghi, Majid Janzamin, and Anima Anandkumar · 2016
Earlier work this paper cites.
L1-regularized neural networks are improperly learnable in polynomial time
Yuchen Zhang, Jason D Lee, and Michael I Jordan · 2016
Earlier work this paper cites.
Generalization and refinement of the integro-local stone theorem for sums of random vectors
Alexandr A Borovkov · 2017
Earlier work this paper cites.
Sgd learns the conjugate kernel class of the network
Amit Daniely · 2017
Earlier work this paper cites.
Reliably learning the relu in polynomial time
Surbhi Goel, Varun Kanade, Adam Klivans, and Justin Thaler · 2017
Earlier work this paper cites.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Earlier work this paper cites.
Learning relus via gradient descent
Mahdi Soltanolkotabi · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Electron-proton dynamics in deep learning
Qiuyi Zhang, Rina Panigrahy, and Sushant Sachdeva · 2017
Earlier work this paper cites.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon · 2017
Earlier work this paper cites.
Universal transformers
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser · 2018
Earlier work this paper cites.
Saber: Module-lwr based key exchange, cpa-secure encryption and cca-secure kem
Jan-Pieter D’Anvers, Angshuman Karmakar, Sujoy Sinha Roy, and Frederik Vercauteren · 2018
Earlier work this paper cites.
Invariance principle on the slice
Yuval Filmus, Guy Kindler, Elchanan Mossel, and Karl Wimmer · 2018
Earlier work this paper cites.
Learning two-layer neural networks with symmetric inputs
Rong Ge, Rohith Kuditipudi, Zhize Li, and Xiang Wang · 2018
Earlier work this paper cites.
Learning one convolutional layer with overlapping patches
Surbhi Goel, Adam R. Klivans, and Raghu Meka · 2018
Earlier work this paper cites.
Learning one-hidden-layer neural networks with landscape design
Rong Ge, Jason D Lee, and Tengyu Ma · 2018
Earlier work this paper cites.
Learning one-hidden-layer neural networks under general input distributions
Weihao Gao, Ashok Vardhan Makkuva, Sewoong Oh, and Pramod Viswanath · 2018
Earlier work this paper cites.
Learning mixtures of linear regressions with nearly optimal complexity
Yuanzhi Li and Yingyu Liang · 2018
Earlier work this paper cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2019
Earlier work this paper cites.
Learning two layer rectified neural networks in polynomial time
Ainesh Bakshi, Rajesh Jayaram, and David P Woodruff · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Learning neural networks with two nonlinear layers in polynomial time
Surbhi Goel and Adam R Klivans · 2019
Cited alongside, same era.
Breaking the gridlock in mixture-of-experts: Consistent and efficient algorithms
Ashok Makkuva, Pramod Viswanath, Sreeram Kannan, and Sewoong Oh · 2019
Cited alongside, same era.
Gradient descent for one-hidden-layer neural networks: Polynomial convergence and sq lower bounds
Santosh Vempala and John Wilmes · 2019
Cited alongside, same era.
Learning one-hidden-layer relu networks via gradient descent
Xiao Zhang, Yaodong Yu, Lingxiao Wang, and Quanquan Gu · 2019
Cited alongside, same era.
Formal language recognition by hard attention transformers: Perspectives from circuit complexity
Yiding Hao, Dana Angluin, and Robert Frank · 2022
Later among the works it cites.
Vision transformers provably learn spatial structure
Samy Jelassi, Michael Eli Sander, and Yuanzhi Li · 2022
Later among the works it cites.
Saturated transformers are constant-depth threshold circuits
William Merrill, Ashish Sabharwal, and Noah A Smith · 2022
Later among the works it cites.
Sgd learning on neural networks: leap complexity and saddle-to-saddle dynamics
Emmanuel Abbe, Enric Boix Adsera, and Theodor Misiakiewicz · 2023
Later among the works it cites.
On learning gaussian multi-index models with gradient flow
Alberto Bietti, Joan Bruna, and Loucas Pillaud-Vivien · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the Ability and Limitations of Transformers to Recognize Formal Languages
Satwik Bhattamishra, Kabir Ahuja, and Navin Goyal · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Learning mixtures of linear regressions in subexponential time via fourier moments
Sitan Chen, Jerry Li, and Zhao Song · 2020
Cited alongside, same era.
Learning polynomials in few relevant dimensions
Sitan Chen and Raghu Meka · 2020
Cited alongside, same era.
Approximation schemes for relu regression
Ilias Diakonikolas, Surbhi Goel, Sushrut Karmalkar, Adam R Klivans, and Mahdi Soltanolkotabi · 2020
Cited alongside, same era.
Small covers for near-zero sets of polynomials and learning latent variable models
Ilias Diakonikolas and Daniel M Kane · 2020
Cited alongside, same era.
Small covers for near-zero sets of polynomials and learning latent variable models
Ilias Diakonikolas and Daniel M Kane · 2020
Cited alongside, same era.
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Later among the works it cites.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection
Yu Bai, Fan Chen, Huan Wang, Caiming Xiong, and Song Mei · 2023
Later among the works it cites.
Learning narrow one-hidden-layer relu networks
Sitan Chen, Zehao Dou, Surbhi Goel, Adam R Klivans, and Raghu Meka · 2023
Later among the works it cites.
What makes a good fisherman? linear regression under self-selection bias
Yeshwanth Cherapanamjeri, Constantinos Daskalakis, Andrew Ilyas, and Manolis Zampetakis · 2023
Later among the works it cites.
Learning polynomial transformations via generalized tensor decompositions
Sitan Chen, Jerry Li, Yuanzhi Li, and Anru R Zhang · 2023
Later among the works it cites.
A faster and simpler algorithm for learning shallow networks
Sitan Chen and Shyam Narayanan · 2023
Later among the works it cites.
Muse: Text-to-image generation via masked generative transformers
Huiwen Chang, Han Zhang, Jarred Barber, AJ Maschinot, Jose Lezama, Lu Jiang, Ming-Hsuan Yang, Kevin Murphy, William T Freeman, Michael Rubinstein, et al · 2023
Later among the works it cites.
On the optimization and generalization of multi-head attention
Puneesh Deora, Rouzbeh Ghaderi, Hossein Taheri, and Christos Thrampoulidis · 2023
Later among the works it cites.
Efficiently learning one-hidden-layer relu networks via schur polynomials
Ilias Diakonikolas and Daniel M Kane · 2023
Later among the works it cites.
Learning two-layer neural networks, one (giant) step at a time
Yatin Dandi, Florent Krzakala, Bruno Loureiro, Luca Pesce, and Ludovic Stephan · 2023
Later among the works it cites.
Computational complexity of learning neural networks: Smoothness and degeneracy
Amit Daniely, Nathan Srebro, and Gal Vardi · 2023
Later among the works it cites.
Tinystories: How small can language models be and still speak coherent english?
Ronen Eldan and Yuanzhi Li · 2023
Later among the works it cites.
What can a single attention layer learn? a study through the random features lens
Hengyu Fu, Tianyu Guo, Yu Bai, and Song Mei · 2023
Later among the works it cites.
The emergence of clusters in self-attention dynamics
Borjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, and Philippe Rigollet · 2023
Later among the works it cites.
A mathematical perspective on transformers
Borjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, and Philippe Rigollet · 2023
Later among the works it cites.
Looped transformers as programmable computers
Angeliki Giannou, Shashank Rajput, Jy-yong Sohn, Kangwook Lee, Jason D Lee, and Dimitris Papailiopoulos · 2023
Later among the works it cites.
Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, et al · 2023
Later among the works it cites.
Transformers learn shortcuts to automata
Bingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang · 2023
Later among the works it cites.
Textbooks are all you need ii: phi-1.5 technical report
Yuanzhi Li, Sébastien Bubeck, Ronen Eldan, Allie Del Giorno, Suriya Gunasekar, and Yin Tat Lee · 2023
Later among the works it cites.
How do transformers learn topic structure: Towards a mechanistic understanding
Yuchen Li, Yuanzhi Li, and Andrej Risteski · 2023
Later among the works it cites.
Hongkang Li, Meng Wang, Sijia Liu, and Pin-Yu Chen · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
On the role of attention in prompt-tuning
Samet Oymak, Ankit Singh Rawat, Mahdi Soltanolkotabi, and Christos Thrampoulidis · 2023
Later among the works it cites.
Representational strengths and limitations of transformers
Clayton Sanford, Daniel Hsu, and Matus Telgarsky · 2023
Later among the works it cites.
Transformers as support vector machines
Davoud Ataee Tarzanagh, Yingcong Li, Christos Thrampoulidis, and Samet Oymak · 2023
Later among the works it cites.
Sequence length independent norm-based generalization bounds for transformers
Jacob Trauger and Ambuj Tewari · 2023
Later among the works it cites.
Scan and snap: Understanding training dynamics and token composition in 1-layer transformer
Yuandong Tian, Yiping Wang, Beidi Chen, and Simon Du · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
Johannes Von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2023
Later among the works it cites.
Trained transformers learn linear models in-context
Ruiqi Zhang, Spencer Frei, and Peter L Bartlett · 2023
Later among the works it cites.
An analysis of attention via the lens of exchangeability and latent variable models, 2023
Yufeng Zhang, Boyi Liu, Qi Cai, Lingxiao Wang, and Zhaoran Wang · 2023
Later among the works it cites.