Fetching the paper…
Reading the bibliography…
Recurrent neural networks are effective models to process sequences.
Language Models are Few-Shot Learners. In NeurIPS , Vol. 33. 1877–1901
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, et al · 1901
Earlier work this paper cites.
“Cloze Procedure”: A New Tool for Measuring Readability
Wilson L. Taylor. 1953 · 1953
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey E. Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
Schemata and Sequential Thought Processes in PDP Models
D. E. Rumelhart, P. Smolensky, J. L. McClelland, and G. E. Hinton. 1986 · 1986
Earlier work this paper cites.
Optimal Brain Damage
Yann LeCun, John S. Denker, and Sara A. Solla. 1990 · 1990
Earlier work this paper cites.
Adaptive Mixtures of Local Experts
R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton. 1991 · 1991
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Chenguang Wang, Zihao Ye, Aston Zhang, Zheng Zhang, and Alexander J. Smola. 2020 · 2002
Earlier work this paper cites.
Longformer: The Long-Document Transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
On the Effect of Dropping Layers of Pre-trained Transformer Models
Hassan Sajjad, Fahim Dalvi, Nadir Durrani, and Preslav Nakov. 2020 · 2004
Earlier work this paper cites.
Synthesizer: Rethinking Self-Attention in Transformer Models
Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng. 2020 · 2005
Earlier work this paper cites.
Linformer: Self-Attention with Linear Complexity
Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma. 2020 · 2006
Earlier work this paper cites.
Challenges and Advances in Parallel Sparse Matrix-Matrix Multiplication. In 2008 37th International Conference on Parallel Processing . 503–510
A. Buluc and J. R. Gilbert. 2008 · 2008
Earlier work this paper cites.
Finding Fast Transformers: One-Shot Neural Architecture Search by Component Composition
Henry Tsai, Jayden Ooi, Chun-Sung Ferng, Hyung Won Chung, and Jason Riesa. 2020 · 2008
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks. In AISTATS , Vol. 9. 249–256
Xavier Glorot and Yoshua Bengio. 2010 · 2010
Earlier work this paper cites.
Transformer-based Online Speech Recognition with Decoder-end Adaptive Computation Steps
Mohan Li, Catalin Zorila, and Rama Doddipatla. 2020 · 2011
Earlier work this paper cites.
Large Text Compression Benchmark
Matt Mahoney. 2011 · 2011
Earlier work this paper cites.
Estimating or Propagating Gradients Through Stochastic Neurons
Yoshua Bengio. 2013b · 2013
Earlier work this paper cites.
Semantic Parsing on Freebase from Question-Answer Pairs. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing . 1533–1544
Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013 · 2013
Earlier work this paper cites.
Saving memory using gradient-checkpointing
OpenAI. 2013 · 2013
Earlier work this paper cites.
Do Deep Nets Really Need to be Deep?. In NIPS , Vol. 27
Jimmy Ba and Rich Caruana. 2014 · 2014
Earlier work this paper cites.
Findings of the 2014 Workshop on Statistical Machine Translation. In SIGMT . 12–58
Ondrej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, et al · 2014
Earlier work this paper cites.
Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation. In EMNLP . 1724–1734
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, et al · 2014
Earlier work this paper cites.
Sequence to Sequence Learning with Neural Networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Earlier work this paper cites.
TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, et al · 2015
Earlier work this paper cites.
Neural Machine Translation by Jointly Learning to Align and Translate. In ICLR
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
NICE: Non-linear Independent Components Estimation. In ICLR
Laurent Dinh, David Krueger, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Earlier work this paper cites.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. In ICML , Vol. 37. 448–456
Sergey Ioffe and Christian Szegedy. 2015 · 2015
Earlier work this paper cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition. In ICASSP . 4960–4964
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals. 2016 · 2016
Earlier work this paper cites.
Training Deep Nets with Sublinear Memory Cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Long Short-Term Memory-Networks for Machine Reading. In EMNLP . 551–561
Jianpeng Cheng, Li Dong, and Mirella Lapata. 2016 · 2016
Earlier work this paper cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville. 2016 · 2016
Earlier work this paper cites.
Adaptive Computation Time for Recurrent Neural Networks
Alex Graves. 2016 · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition. In CVPR . 770–778
K. He, X. Zhang, S. Ren, and J. Sun. 2016 · 2016
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text. In EMNLP . 2383–2392
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Density estimation using Real NVP. In ICLR
Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio. 2017 · 2017
Earlier work this paper cites.
The Reversible Residual Network: Backpropagation Without Storing Activations. In NeurIPS , Vol. 30
Aidan N Gomez, Mengye Ren, Raquel Urtasun, and Roger B Grosse. 2017 · 2017
Earlier work this paper cites.
GPU Kernels for Block-Sparse Weights
Scott Gray, Alec Radford, and Diederik P. Kingma. 2017 · 2017
Earlier work this paper cites.
The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables. In ICLR
Chris J. Maddison, Andriy Mnih, and Yee Whye Teh. 2017 · 2017
Earlier work this paper cites.
Pointer Sentinel Mixture Models. In ICLR
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017 · 2017
Earlier work this paper cites.
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.. In ICLR
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, et al · 2017
Earlier work this paper cites.
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, et al · 2017
Earlier work this paper cites.
Neural Architecture Search with Reinforcement Learning. In ICLR
Barret Zoph and Quoc V. Le. 2017 · 2017
Cited alongside, same era.
World gross electricity production, by source, 2018
IEA. 2018 · 2018
Cited alongside, same era.
Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference. In IEEE/CVF . 2704–2713
B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, et al · 2018
Cited alongside, same era.
Sharp Nearby, Fuzzy Far Away: How Neural Language Models Use Context. In ACL . 284–294
Urvashi Khandelwal, He He, Peng Qi, and Dan Jurafsky. 2018 · 2018
Cited alongside, same era.
Mixed Precision Training. In ICLR
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, et al · 2018
Cited alongside, same era.
ListOps: A Diagnostic Dataset for Latent Tree Learning. In NAACL
Depth-Adaptive Transformer. In ICLR
Maha Elbayad, Jiatao Gu, Edouard Grave, and Michael Auli. 2020 · 2020
Later among the works it cites.
Attention in Natural Language Processing
Andrea Galassi, Marco Lippi, and Paolo Torroni. 2020 · 2020
Later among the works it cites.
Conformer: Convolution-augmented Transformer for Speech Recognition. In Interspeech . 5036–5040
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, et al · 2020
Later among the works it cites.
Sara Hooker. 2020 · 2020
Later among the works it cites.
Improving Transformer Optimization Through Better Initialization. In ICML , Vol. 119. 4475–4483
Xiao Shi Huang, Felipe Pérez, Jimmy Ba, and Maksims Volkovs. 2020 · 2020
Later among the works it cites.
TinyBERT: Distilling BERT for Natural Language Understanding. In EMNLP . 4163–4174
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nikita Nangia and Samuel R. Bowman. 2018 · 2018
Cited alongside, same era.
Don’t Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization. In EMNLP . 1797–1807
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018 · 2018
Cited alongside, same era.
Scaling Neural Machine Translation. In Proceedings of the Third Conference on Machine Translation: Research Papers . 1–9
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli. 2018 · 2018
Cited alongside, same era.
Efficient Neural Architecture Search via Parameter Sharing. In ICML , Vol. 80. 4092–4101
Hieu Pham, Melody Y. Guan, Barret Zoph, Quoc V. Le, and Jeff Dean. 2018 · 2018
Cited alongside, same era.
Improving Language Understanding by Generative Pre-Training
Alec Radford and Karthik Narasimhan. 2018 · 2018
Cited alongside, same era.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2018 · 2018
Cited alongside, same era.
How Does Batch Normalization Help Optimization?. In NeurIPS , Vol. 31
Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry. 2018 · 2018
Cited alongside, same era.
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, et al · 2020
Later among the works it cites.
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention. In ICML , Vol. 119. 5156–5165
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. 2020 · 2020
Later among the works it cites.
Reformer: The Efficient Transformer. In ICLR
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. 2020 · 2020
Later among the works it cites.
Big Transfer (BiT): General Visual Representation Learning. In ECCV , Vol. 12350. 491–507
Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, et al · 2020
Later among the works it cites.
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations. In ICLR
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 2020
Later among the works it cites.
Reconciling Modern Deep Learning with Traditional Optimization Analyses: The Intrinsic Learning Rate. In NeurIPS , Vol. 33. 14544–14555
Zhiyuan Li, Kaifeng Lyu, and Sanjeev Arora. 2020 · 2020
Later among the works it cites.
Estimating Carbon Emissions of Artificial Intelligence [Opinion]
Alexandra Luccioni, Alexandre Lacoste, and Victor Schmidt. 2020 · 2020
Later among the works it cites.
Specaugment on Large Scale Datasets. In ICASSP . 6879–6883
Daniel S. Park, Yu Zhang, Chung-Cheng Chiu, Youzheng Chen, Bo Li, William Chan, et al · 2020
Later among the works it cites.
When BERT Plays the Lottery, All Tickets Are Winning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) . 3208–3229
Sai Prasanna, Anna Rogers, and Anna Rumshisky. 2020 · 2020
Later among the works it cites.
Blockwise Self-Attention for Long Document Understanding. In EMNLP . 2555–2565
Jiezhong Qiu, Hao Ma, Omer Levy, Wen-tau Yih, Sinong Wang, and Jie Tang. 2020 · 2020
Later among the works it cites.
Compressive Transformers for Long-Range Sequence Modelling. In ICLR
Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier, and Timothy P. Lillicrap. 2020 · 2020
Later among the works it cites.
Energy and Policy Considerations for Modern Deep Learning Research
Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2020 · 2020
Later among the works it cites.
Efficient Transformers: A Survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2020b · 2020
Later among the works it cites.
Fast Transformers with Clustered Attention. In NeurIPS
Apoorv Vyas, Angelos Katharopoulos, and François Fleuret. 2020 · 2020
Later among the works it cites.
Benchmarking the Performance and Energy Efficiency of AI Accelerators for AI Training. In CCGRID . 744–751
Yuxin Wang, Qiang Wang, Shaohuai Shi, Xin He, Zhenheng Tang, Kaiyong Zhao, et al · 2020
Later among the works it cites.
Transformers: State-of-the-Art Natural Language Processing. In EMNLP . 38–45
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, et al · 2020
Later among the works it cites.
Lite Transformer with Long-Short Range Attention. In ICLR
Zhanghao Wu, Zhijian Liu, Ji Lin, Yujun Lin, and Song Han. 2020 · 2020
Later among the works it cites.
Self-Training With Noisy Student Improves ImageNet Classification.. In CVPR . 10684–10695
Qizhe Xie, Minh-Thang Luong, Eduard H. Hovy, and Quoc V. Le. 2020 · 2020
Later among the works it cites.
On Layer Normalization in the Transformer Architecture. In ICML , Vol. 119. 10524–10533
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, et al · 2020
Later among the works it cites.
Big Bird: Transformers for Longer Sequences. In NeurIPS , Vol. 33. 17283–17297
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, et al · 2020
Later among the works it cites.
Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss. In ICASSP . 7829–7833
Qian Zhang, Han Lu, Hasim Sak, Anshuman Tripathi, Erik McDermott, Stephen Koo, et al · 2020
Later among the works it cites.
LambdaNetworks: Modeling long-range Interactions without Attention. In ICLR
Irwan Bello. 2021 · 2021
Closest in time.
Rethinking Attention with Performers. In ICLR
Krzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, et al · 2021
Closest in time.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In ICLR
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, et al · 2021
Closest in time.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
William Fedus, Barret Zoph, and Noam Shazeer. 2021 · 2021
Closest in time.
On Improving Deep Learning Trace Analysis with System Call Arguments. In MSR . 120–130
Quentin Fournier, Daniel Aloise, Seyed Vahid Azhari, and François Tetreault. 2021 · 2021
Closest in time.
Transformers in Vision: A Survey
Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, and Mubarak Shah. 2021 · 2021
Closest in time.
Transformers with Competitive Ensembles of Independent Mechanisms
Alex Lamb, Di He, Anirudh Goyal, Guolin Ke, Chien-Feng Liao, Mirco Ravanelli, et al · 2021
Closest in time.
Tianyang Lin, Yuxin Wang, Xiangyang Liu, and Xipeng Qiu. 2021 · 2021
Closest in time.
Hanxiao Liu, Zihang Dai, David R. So, and Quoc V. Le. 2021a · 2021
Closest in time.
Do Transformer Modifications Transfer Across Implementations and Applications?
Sharan Narang, Hyung Won Chung, Yi Tay, William Fedus, Thibault Fevry, Michael Matena, et al · 2021
Closest in time.
Efficient Content-Based Sparse Attention with Routing Transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier. 2021 · 2021
Closest in time.
SparseBERT: Rethinking the Importance Analysis in Self-attention. In ICML , Vol. 139. 9547–9557
Han Shi, Jiahui Gao, Xiaozhe Ren, Hang Xu, Xiaodan Liang, Zhenguo Li, et al · 2021
Closest in time.
Training with Quantization Noise for Extreme Model Compression. In ICLR
Pierre Stock, Angela Fan, Benjamin Graham, Edouard Grave, Rémi Gribonval, Herve Jegou, et al · 2021
Closest in time.
Attention Is All You Need In Speech Separation. In ICASSP . 21–25
Cem Subakan, Mirco Ravanelli, Samuele Cornell, Mirko Bronzi, and Jianyuan Zhong. 2021 · 2021
Closest in time.
Long Range Arena : A Benchmark for Efficient Transformers. In ICLR
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, et al · 2021
Closest in time.
MLP-Mixer: An all-MLP Architecture for Vision
Ilya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, et al · 2021
Closest in time.
Nyströmformer: A Nyström-Based Algorithm for Approximating Self-Attention
Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, et al · 2021
Closest in time.
MetaFormer Is Actually What You Need for Vision. In CVPR . 10819–10829
Weihao Yu, Mi Luo, Pan Zhou, Chenyang Si, Yichen Zhou, Xinchao Wang, Jiashi Feng, and Shuicheng Yan. 2022 · 2022
Closest in time.