Fetching the paper…
Reading the bibliography…
Attention plays a fundamental role in both natural and artificial intelligence systems.
The capacity of feedforward neural networks
Pierre Baldi and Roman Vershynin · 1901
Earlier work this paper cites.
Neural networks, orientations of the hypercube and algebraic threshold functions
P. Baldi · 1988
Earlier work this paper cites.
A back-propagation programmed network that simulates response properties of a subset of posterior parietal neurons
David Zipser and Richard A Andersen · 1988
Earlier work this paper cites.
Asymptotics of the logarithm of the number of threshold functions of the algebra of logic
Yu A Zuev · 1989
Earlier work this paper cites.
Combinatorial-probability and geometric methods in threshold logic
Yu A Zuev · 1991
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to dynamic recurrent networks
Jürgen Schmidhuber · 1992
Earlier work this paper cites.
On the probability that a random ± \pm 1-matrix is singular
Jeff Kahn, János Komlós, and Endre Szemerédi · 1995
Earlier work this paper cites.
Emergence of simple-cell receptive field properties by learning a sparse code for natural images
Bruno A Olshausen and David J Field · 1996
Earlier work this paper cites.
Discrete mathematics of neural networks: selected topics
Martin Anthony · 2001
Earlier work this paper cites.
Neurobiology of attention
Laurent Itti, Geraint Rees, and John K Tsotsos · 2005
Earlier work this paper cites.
Neurobiology of attention regulation and its disorders
Amy F.T. Arnsten and Francisco X. Castellanos · 2010
Earlier work this paper cites.
Permutationless many-jet event reconstruction with symmetry preserving attention networks
M. Fenton, A. Shmakov, T. Ho, S. Hsu, D. Whiteson, and P. Baldi · 2010
Cited alongside, same era.
Cognitive neuroscience of attention
Michael I Posner · 2011
Cited alongside, same era.
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyung Hyun Cho, and Yoshua Bengio · 2015
Cited alongside, same era.
Attention-based models for speech recognition
Jan Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio · 2015
Cited alongside, same era.
Neural networks capacity
Pierre Baldi and Roman Vershynin · 2018
Later among the works it cites.
On neuronal capacity
Pierre Baldi and Roman Vershynin · 2018
Later among the works it cites.
Polynomial threshold functions, hyperplane arrangements, and random tensors
Pierre Baldi and Roman Vershynin · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Later among the works it cites.
Set transformer: A framework for attention-based permutation-invariant neural networks
Juho Lee, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Seungjin Choi, and Yee Whye Teh · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Minh-Thang Luong, Hieu Pham, and Christopher D Manning · 2015
Cited alongside, same era.
Using fast weights to attend to the recent past
Jimmy Ba, Geoffrey E Hinton, Volodymyr Mnih, Joel Z Leibo, and Catalin Ionescu · 2016
Cited alongside, same era.
Parameterized neural networks for high-energy physics
Pierre Baldi, Kyle Cranmer, Taylor Faucett, Peter Sadowski, and Daniel Whiteson · 2016
Cited alongside, same era.
Using goal-driven deep learning models to understand sensory cortex
Daniel LK Yamins and James J DiCarlo · 2016
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Later among the works it cites.
Deep Learning in Science
P. Baldi · 2021
Later among the works it cites.
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas · 2021
Later among the works it cites.
Hanxiao Liu, Zihang Dai, David R So, and Quoc V Le · 2021
Later among the works it cites.
Splash: Learnable activation functions for improving accuracy and adversarial robustness
Mohammadamin Tavakoli, Forest Agostinelli, and Pierre Baldi · 2021
Later among the works it cites.