Fetching the paper…
Reading the bibliography…
Transformers have proved effective in many NLP tasks.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019 · 1910
Earlier work this paper cites.
Muse: Parallel multi-scale attention for sequence to sequence learning
Guangxiang Zhao, Xu Sun, Jingjing Xu, Zhiyuan Zhang, and Liangchen Luo. 2019 · 1911
Earlier work this paper cites.
Why adam beats sgd for attention models
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank J. Reddi, Surinder Kumar, and Suvrit Sra. 2019b · 1912
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Y. Simard, and Paolo Frasconi. 1994 · 1994
Earlier work this paper cites.
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shu xin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Li-Wei Wang, and Tie-Yan Liu. 2019 · 2002
Earlier work this paper cites.
Rezero is all you need: Fast convergence at large depth
Thomas C. Bachlechner, Bodhisattwa Prasad Majumder, Huanru Henry Mao, Garrison W. Cottrell, and Julian J. McAuley. 2020 · 2003
Earlier work this paper cites.
Very deep transformers for neural machine translation
Xiaodong Liu, Kevin Duh, Liyuan Liu, and Jianfeng Gao. 2020b · 2008
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio. 2010 · 2010
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli. 2013 · 2013
Earlier work this paper cites.
Findings of the 2014 workshop on statistical machine translation
Ondřej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, et al. 2014 · 2014
Earlier work this paper cites.
Report on the 11th iwslt evaluation campaign, iwslt 2014
Mauro Cettolo, Jan Niehues, Sebastian Stüker, Luisa Bentivogli, and Marcello Federico. 2014 · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015 · 2015
Cited alongside, same era.
All you need is a good init
Dmytro Mishkin and Juan E. Sala Matas. 2015 · 2015
Cited alongside, same era.
Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016 · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016 · 2016
Cited alongside, same era.
Training tips for the transformer model
Martin Popel and Ondrej Bojar. 2018 · 2018
Later among the works it cites.
Graph attention networks
Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. 2018 · 2018
Later among the works it cites.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli. 2019 · 2019
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Music transformer: Generating music with long-term structure
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Ian Simon, Curtis Hawthorne, Noam Shazeer, Andrew M. Dai, Matthew D. Hoffman, Monica Dinculescu, and Douglas Eck. 2019 · 2019
Later among the works it cites.
Transformers without tears: Improving the normalization of self-attention
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Priya Goyal, Piotr Dollár, Ross B. Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
The best of both worlds: Combining recent advances in neural machine translation
Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George Foster, Llion Jones, Niki Parmar, Michael Schuster, Zhi-Feng Chen, Yonghui Wu, and Macduff Hughes. 2018 · 2018
Cited alongside, same era.
How to start training: The effect of initialization and architecture
Boris Hanin and David Rolnick. 2018 · 2018
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, and Tom Goldstein. 2018 · 2018
Cited alongside, same era.
Image transformer
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran. 2018 · 2018
Cited alongside, same era.
The shattered gradients problem: If resnets are the answer, then what is the question?
David Balduzzi, Marcus Frean, Lennox Leary, J. P. Lewis, Kurt Wan-Duo Ma, and Brian McWilliams. 2017a
Cited in the paper.
Toan Q. Nguyen and Julian Salazar. 2019 · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Later among the works it cites.
Stand-alone self-attention in vision models
Prajit Ramachandran, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jonathon Shlens. 2019 · 2019
Later among the works it cites.
Learning deep transformer models for machine translation
Qiang Wang, Bei Li, Tong Xiao, Jingbo Zhu, Changliang Li, Derek F. Wong, and Lidia S. Chao. 2019 · 2019
Later among the works it cites.
Towards adaptive residual network training: A neural-ode perspective
Chengyu Dong, Liyuan Liu, Zichao Li, and Jingbo Shang. 2020 · 2020
Closest in time.
Understanding and improving transformer from a multi-particle dynamic system point of view
Yiping Lu, Zhuohan Li, Di He, Zhiqing Sun, Bin Dong, Tao Qin, Liwei Wang, and Tie-Yan Liu. 2020 · 2020
Closest in time.