Fetching the paper…
Reading the bibliography…
Self-attention has recently been adopted for a wide range of sequence modeling problems.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 1904
Earlier work this paper cites.
Convergence properties of the k-means algorithms
Leon Bottou and Yoshua Bengio. 1995 · 1995
Earlier work this paper cites.
Frequency-sensitive competitive learning for scalable balanced clustering on high-dimensional hyperspheres
Arindam Banerjee and Joydeep Ghosh. 2004 · 2004
Earlier work this paper cites.
Sparse gpu kernels for deep learning
Trevor Gale, Matei Zaharia, Cliff Young, and Erich Elsen. 2020 · 2006
Earlier work this paper cites.
Large text compression benchmark
Matt Mahoney. 2011 · 2011
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville. 2013 · 2013
Earlier work this paper cites.
Learning factored representations in a deep mixture of experts
David Eigen, Marc’Aurelio Ranzato, and Ilya Sutskever. 2013 · 2013
Earlier work this paper cites.
Kyunghyun Cho and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Learning phrase representations using rnn encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Deep sequential neural network
Ludovic Denoyer and Patrick Gallinari. 2014 · 2014
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka. 2014 · 2014
Earlier work this paper cites.
Balanced k-means for clustering
Mikko I Malinen and Pasi Fränti. 2014 · 2014
Earlier work this paper cites.
Clustering is efficient for approximate maximum inner product search
Alex Auvolat, Sarath Chandar, Pascal Vincent, Hugo Larochelle, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Attention-based models for speech recognition
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
DRAW: A recurrent neural network for image generation
Karol Gregor, Ivo Danihelka, Alex Graves, Danilo Jimenez Rezende, and Daan Wierstra. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D Manning. 2015 · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016 · 2016
Earlier work this paper cites.
An online sequence-to-sequence model using partial conditioning
Navdeep Jaitly, Quoc V Le, Oriol Vinyals, Ilya Sutskever, David Sussillo, and Samy Bengio. 2016 · 2016
Cited alongside, same era.
Conditional image generation with pixelcnn decoders
Aaron Van den Oord, Nal Kalchbrenner, Lasse Espeholt, Oriol Vinyals, Alex Graves, et al. 2016 · 2016
Cited alongside, same era.
Scaling memory-augmented neural networks with sparse reads and writes
Jack Rae, Jonathan J Hunt, Ivo Danihelka, Timothy Harley, Andrew W Senior, Gregory Wayne, Alex Graves, and Timothy Lillicrap. 2016 · 2016
Cited alongside, same era.
Improving neural language models with a continuous cache
Edouard Grave, Armand Joulin, and Nicolas Usunier. 2017 · 2017
Cited alongside, same era.
Learning what’s easy: Fully differentiable neural easy-first taggers
André F. T. Martins and Julia Kreutzer. 2017 · 2017
Cited alongside, same era.
Pointer sentinel mixture models
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2018
Later among the works it cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern. 2018 · 2018
Later among the works it cites.
Character-level language modeling with deeper self-attention
Rami Al-Rfou, Dokook Choe, Noah Constant, Mandy Guo, and Llion Jones. 2019 · 2019
Later among the works it cites.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli. 2019 · 2019
Later among the works it cites.
Learning classifiers with fenchel-young losses: Generalized entropies, margins, and algorithms
Mathieu Blondel, André F. T. Martins, and Vlad Niculae. 2019 · 2019
Later among the works it cites.
Adaptively sparse transformers
Gonçalo M Correia, Vlad Niculae, and André FT Martins. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, and Jeff Dean. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Pixelsnail: An improved autoregressive generative model
Xi Chen, Nikhil Mishra, Mostafa Rohaninejad, and Pieter Abbeel. 2018 · 2018
Cited alongside, same era.
Monotonic chunkwise attention
Chung-Cheng Chiu* and Colin Raffel*. 2018 · 2018
Cited alongside, same era.
Music transformer: Generating music with long-term structure
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Ian Simon, Curtis Hawthorne, Noam Shazeer, Andrew M Dai, Matthew D Hoffman, Monica Dinculescu, and Douglas Eck. 2018 · 2018
Cited alongside, same era.
Glow: Generative flow with invertible 1x1 convolutions
Durk P Kingma and Prafulla Dhariwal. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019 · 2019
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Look harder: A neural machine translation model with hard attention
Sathish Reddy Indurthi, Insoo Chung, and Sangha Kim. 2019 · 2019
Later among the works it cites.
Large memory layers with product keys
Guillaume Lample, Alexandre Sablayrolles, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2019 · 2019
Later among the works it cites.
Multi-task deep neural networks for natural language understanding
Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao. 2019 · 2019
Later among the works it cites.
Adaptive attention span in transformers
Sainbayar Sukhbaatar, Édouard Grave, Piotr Bojanowski, and Armand Joulin. 2019 · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
Later among the works it cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Closest in time.
Reformer: The efficient transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. 2020 · 2020
Closest in time.
Do transformers need deep long-range memory?
Jack Rae and Ali Razavi. 2020 · 2020
Closest in time.
Compressive transformers for long-range sequence modelling
Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier, and Timothy P. Lillicrap. 2020 · 2020
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron C. Courville, Ruslan Salakhutdinov, Richard S. Zemel, and Yoshua Bengio. 2015 · 2057
Closest in time.