Fetching the paper…
Reading the bibliography…
Despite the widespread adoption of Transformer models for NLP tasks, the expressive power of these models is not well-understood.
Approximation by superpositions of a sigmoidal function
George Cybenko · 1989
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
Kurt Hornik · 1991
Earlier work this paper cites.
The fast Johnson–Lindenstrauss transform and approximate nearest neighbors
Nir Ailon and Bernard Chazelle · 2009
Earlier work this paper cites.
Learning binary codes for high-dimensional data using bilinear projections
Yunchao Gong, Sanjiv Kumar, Henry A Rowley, and Svetlana Lazebnik · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean · 2013
Earlier work this paper cites.
Rigid-motion scattering for image classification
Laurent Sifre and Stéphane Mallat · 2014
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D. Manning · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Xception: Deep learning with depthwise separable convolutions
François Chollet · 2017
Earlier work this paper cites.
Approximating continuous functions by relu nets of minimal width
Boris Hanin and Mark Sellke · 2017
Cited alongside, same era.
Depthwise separable convolutions for neural machine translation
Lukasz Kaiser, Aidan N Gomez, and Francois Chollet · 2017
Cited alongside, same era.
The expressive power of neural networks: A view from the width
Zhou Lu, Hongming Pu, Feicheng Wang, Zhiqiang Hu, and Liwei Wang · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser · 2018
Cited alongside, same era.
Visualizing and measuring the geometry of BERT
Andy Coenen, Emily Reif, Ann Yuan, Been Kim, Adam Pearce, Fernanda Viégas, and Martin Wattenberg · 2019
Closest in time.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D Manning · 2019
Closest in time.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Closest in time.
On the Turing completeness of modern neural network architectures
Jorge Pérez, Javier Marinković, and Pablo Barceló · 2019
Closest in time.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
ResNet with one-neuron hidden layers is a universal approximator
Hongzhou Lin and Stefanie Jegelka · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Cited alongside, same era.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning · 2019
Cited alongside, same era.
Closest in time.
Universal approximations of permutation invariant/equivariant functions by deep neural networks
Akiyoshi Sannai, Yuuki Takai, and Matthieu Cordonnier · 2019
Closest in time.
Analyzing the structure of attention in a transformer language model
Jesse Vig and Yonatan Belinkov · 2019
Closest in time.
Pay less attention with lightweight and dynamic convolutions
Felix Wu, Angela Fan, Alexei Baevski, Yann N Dauphin, and Michael Auli · 2019
Closest in time.
XLNet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le · 2019
Closest in time.