Fetching the paper…
Reading the bibliography…
Attention is a powerful component of modern neural networks across a wide variety of domains.
On estimating regression
Elizbar A Nadaraya · 1964
Earlier work this paper cites.
Optimal plans for dynamic programming problems
CJ Himmelberg, T Parthasarathy, and FS VanVleck · 1976
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
Feynman-kac formulae: Genealogical and interacting particle systems with applications, probability and its applications, 2004
P Del Moral · 2004
Earlier work this paper cites.
Large population stochastic dynamic games: closed-loop mckean-vlasov systems and the nash certainty equivalence principle
Minyi Huang, Roland P Malhamé, Peter E Caines, et al · 2006
Earlier work this paper cites.
Jeux à champ moyen. i–le cas stationnaire
Jean-Michel Lasry and Pierre-Louis Lions · 2006
Earlier work this paper cites.
Non-linear markov chain monte carlo
Christophe Andrieu, Ajay Jasra, Arnaud Doucet, and Pierre Del Moral · 2007
Earlier work this paper cites.
Optimal Transport: Old and New
C. Villani · 2008
Earlier work this paper cites.
Probabilistic graphical models: principles and techniques
Daphne Koller and Nir Friedman · 2009
Earlier work this paper cites.
Infinite Dimensional Analysis: A Hitchhiker’s Guide
C.D. Aliprantis and K.C. Border · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
One-dimensional empirical measures, order statistics and kantorovich transport distances
Sergey Bobkov and Michel Ledoux · 2014
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Cited alongside, same era.
Rényi divergence and kullback-leibler divergence
Tim Van Erven and Peter Harremos · 2014
Cited alongside, same era.
Optimal Transport for Applied Mathematicians: Calculus of Variations, PDEs, and Modeling
F. Santambrogio · 2015
Cited alongside, same era.
End-to-end memory networks
Sainbayar Sukhbaatar, Jason Weston, Rob Fergus, et al · 2015
Cited alongside, same era.
On the expressive power of deep learning: A tensor analysis
Nadav Cohen, Or Sharir, and Amnon Shashua · 2016
Cited alongside, same era.
Understanding deep convolutional networks
Stéphane Mallat · 2016
Cited alongside, same era.
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Łukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Later among the works it cites.
Emergent tool use from multi-agent autocurricula
Bowen Baker, Ingmar Kanitscheider, Todor Markov, Yi Wu, Glenn Powell, Bob McGrew, and Igor Mordatch · 2019
Later among the works it cites.
What does bert look at? an analysis of bert’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning · 2019
Later among the works it cites.
Negated and misprimed probes for pretrained language models: Birds can talk, but cannot fly, 2019
Nora Kassner and Hinrich Schütze · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bidirectional attention flow for machine comprehension
Minjoon Seo, Aniruddha Kembhavi, Ali Farhadi, and Hannaneh Hajishirzi · 2016
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio · 2017
Cited alongside, same era.
Smooth regression analysis
Geoffrey S Watson · 2017
Cited alongside, same era.
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Later among the works it cites.
Understanding and improving transformer from a multi-particle dynamic system point of view
Yiping Lu, Zhuohan Li, Di He, Zhiqing Sun, Bin Dong, Tao Qin, Liwei Wang, and Tie-Yan Liu · 2019
Later among the works it cites.
Attention in deep learning, 2019
Alex Smola and Aston Zhang · 2019
Later among the works it cites.
Bert rediscovers the classical nlp pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick · 2019
Later among the works it cites.
On the computational power of transformers and its implications in sequence modeling
Satwik Bhattamishra, Arkil Patel, and Navin Goyal · 2020
Closest in time.
Infinite attention: Nngp and ntk for deep attention networks
Jiri Hron, Yasaman Bahri, Jascha Sohl-Dickstein, and Roman Novak · 2020
Closest in time.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Closest in time.
The lipschitz constant of self-attention
Hyunjik Kim, George Papamakarios, and Andriy Mnih · 2020
Closest in time.
Limits to depth efficiencies of self-attention
Yoav Levine, Noam Wies, Or Sharir, Hofit Bata, and Amnon Shashua · 2020
Closest in time.