Fetching the paper…
Reading the bibliography…
Understanding the fundamental mechanism behind the success of transformer networks is still an open problem in the deep learning literature.
Attention interpretability across nlp tasks, 2019
Shikhar Vashishth, Shyam Upadhyay, Gaurav Singh Tomar, and Manaal Faruqui · 1909
Earlier work this paper cites.
On general minimax theorems
Maurice Sion · 1958
Earlier work this paper cites.
Adaptive regression and model selection in data mining problems
Sergey Bakin et al · 1999
Earlier work this paper cites.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Model selection and estimation in regression with grouped variables
Ming Yuan and Yi Lin · 2006
Earlier work this paper cites.
L1 regularization in infinite dimensional feature spaces
Saharon Rosset, Grzegorz Swirszcz, Nathan Srebro, and Ji Zhu · 2007
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R Bowman · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
How do infinite width bounded norm networks look in function space?
Pedro Savarese, Itay Evron, Daniel Soudry, and Nathan Srebro · 2019
Cited alongside, same era.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov · 2019
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Cited alongside, same era.
Fourier neural operator for parametric partial differential equations
Zongyi Li, Nikola Kovachki, Kamyar Azizzadenesheli, Burigede Liu, Kaushik Bhattacharya, Andrew Stuart, and Anima Anandkumar · 2020
Cited alongside, same era.
Neural networks are convex regularizers: Exact polynomial-time convex optimization formulations for two-layer networks
Raftmlp: Do mlp-based models dream of winning over computer vision?
Yuki Tatsunami and Masato Taki · 2021
Later among the works it cites.
Mlp-mixer: An all-mlp architecture for vision
Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, et al · 2021
Later among the works it cites.
Resmlp: Feedforward networks for image classification with data-efficient training
Hugo Touvron, Piotr Bojanowski, Mathilde Caron, Matthieu Cord, Alaaeldin El-Nouby, Edouard Grave, Gautier Izacard, Armand Joulin, Gabriel Synnaeve, Jakob Verbeek, et al · 2021
Later among the works it cites.
Metaformer is actually what you need for vision
Weihao Yu, Mi Luo, Pan Zhou, Chenyang Si, Yichen Zhou, Xinchao Wang, Jiashi Feng, and Shuicheng Yan · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mert Pilanci and Tolga Ergen · 2020
Cited alongside, same era.
Attention is not all you need: Pure attention loses rank doubly exponentially with depth, 2021
Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas · 2021
Cited alongside, same era.
Convex geometry and duality of over-parameterized neural networks
Tolga Ergen and Mert Pilanci · 2021
Cited alongside, same era.
Is attention better than matrix decomposition?
Zhengyang Geng, Meng-Hao Guo, Hongxu Chen, Xia Li, Ke Wei, and Zhouchen Lin · 2021
Cited alongside, same era.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy · 2021
Cited alongside, same era.
Adaptive fourier neural operators: Efficient token mixers for transformers
John Guibas, Morteza Mardani, Zongyi Li, Andrew Tao, Anima Anandkumar, and Bryan Catanzaro · 2021
Cited alongside, same era.
Fnet: Mixing tokens with fourier transforms
James Lee-Thorp, Joshua Ainslie, Ilya Eckstein, and Santiago Ontanon · 2021
Cited alongside, same era.
Global filter networks for image classification
Yongming Rao, Wenliang Zhao, Zheng Zhu, Jiwen Lu, and Jie Zhou · 2021
Cited alongside, same era.
Boaz Barak, Benjamin L Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang · 2022
Closest in time.
Locating and editing factual associations in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Closest in time.
Aaron Mishkin, Arda Sahiner, and Mert Pilanci · 2022
Closest in time.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Closest in time.
Unraveling attention via convex duality: Analysis and interpretations of vision transformers
Arda Sahiner, Tolga Ergen, Batu Ozturkler, John Pauly, Morteza Mardani, and Mert Pilanci · 2022
Closest in time.
On layer normalizations and residual connections in transformers
Sho Takase, Shun Kiyono, Sosuke Kobayashi, and Jun Suzuki · 2022
Closest in time.
The slingshot mechanism: An empirical study of adaptive optimizers and the grokking phenomenon
Vimal Thilak, Etai Littwin, Shuangfei Zhai, Omid Saremi, Roni Paiss, and Joshua Susskind · 2022
Closest in time.