Markov Chains
J. R. Norris · 1997
Earlier work this paper cites.
Elements of information theory
Thomas M Cover and Joy A Thomas · 2006
Earlier work this paper cites.
Matrix analysis
Roger A Horn and Charles R Johnson · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Benefits of depth in neural networks
Matus Telgarsky · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding, 2018
Original
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford and Karthik Narasimhan · 2018
Earlier work this paper cites.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Earlier work this paper cites.
Are transformers universal approximators of sequence-to-sequence functions?
Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank Reddi, and Sanjiv Kumar · 2020
Earlier work this paper cites.
Rethinking embedding coupling in pre-trained language models
Hyung Won Chung, Thibault Fevry, Henry Tsai, Melvin Johnson, and Sebastian Ruder · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Earlier work this paper cites.
Attention is Turing-complete
Jorge Pérez, Pablo Barceló, and Javier Marinkovic · 2021
Earlier work this paper cites.
Approximating how single head attention learns, 2021
Original
Charlie Snell, Ruiqi Zhong, Dan Klein, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Thinking like transformers
Gail Weiss, Yoav Goldberg, and Eran Yahav · 2021
Earlier work this paper cites.