Fetching the paper…
Reading the bibliography…
This document aims to be a self-contained, mathematically precise overview of transformer architectures and algorithms (*not* results).
A new algorithm for data compression
Philip Gage · 1994
Earlier work this paper cites.
Playing Atari with Deep Reinforcement Learning, December 2013
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
On the number of linear regions of deep neural networks
Guido F Montufar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
A Brief Overview of Deep Learning
Ilya Sutskever · 2015
Earlier work this paper cites.
An overview of gradient descent optimization algorithms
Sebastian Ruder · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
The Illustrated Transformer
Jay Alammar · 2018
Earlier work this paper cites.
Automatic Differentiation in Machine Learning: A Survey
Atilim Gunes Baydin, Barak A. Pearlmutter, Alexey Andreyevich Radul, and Jeffrey Mark Siskind · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S. Sutton, Andrew G. Barto, and Francis Bach · 2018
Cited alongside, same era.
The Illustrated GPT-2 (Visualizing Transformer Language Models)
Jay Alammar · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, et al · 2020
Cited alongside, same era.
A survey of data augmentation approaches for NLP
Steven Y Feng, Varun Gangal, Jason Wei, Sarath Chandar, Soroush Vosoughi, Teruko Mitamura, and Eduard Hovy · 2021
Later among the works it cites.
Master Positional Encoding: Part I
Jonathan Kernes · 2021
Later among the works it cites.
Data preprocessing in NLP
Chris Lemke · 2021
Later among the works it cites.
A Survey of Transformers, June 2021
Tianyang Lin, Yuxin Wang, Xiangyang Liu, and Xipeng Qiu · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
Jack W. Rae, Sebastian Borgeaud, Trevor Cai, et al · 2021
Later among the works it cites.
A rapid and efficient learning rule for biological neural circuits
Eren Sezener, Agnieszka Grabska-Barwińska, Dimitar Kostadinov, Maxime Beau, Sanjukta Krishnagopal, David Budden, Marcus Hutter, Joel Veness, Matthew Botvinick, Claudia Clopath, Michael Häusser, and Peter E. Latham · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A survey of regularization strategies for deep models
Reza Moradi, Reza Berangi, and Behrouz Minaei · 2020
Cited alongside, same era.
NNCP v2: Lossless Data Compression with Transformer
Fabrice Bellard · 2021
Cited alongside, same era.
Inductive Biases and Variable Creation in Self-Attention Mechanisms
Benjamin L. Edelman, Surbhi Goel, Sham Kakade, and Cyril Zhang · 2021
Cited alongside, same era.
A Mathematical Framework for Transformer Circuits
Nelson Elhage · 2021
Cited alongside, same era.
Provable RL with Exogenous Distractors via Multistep Inverse Dynamics, March 2021
Yonathan Efroni, Dipendra Misra, Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2021
Cited alongside, same era.
Later among the works it cites.
Solving Quantitative Reasoning Problems with Language Models
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra · 2022
Closest in time.
Formal algorithms for transformers
M. Phuong and M. Hutter · 2022
Closest in time.
Scott Reed, Konrad Żołna, Emilio Parisotto, et al · 2022
Closest in time.
A comprehensive survey on regularization strategies in machine learning
Yingjie Tian and Yuqi Zhang · 2022
Closest in time.