Fetching the paper…
Reading the bibliography…
Transformer architecture achieves great success in abundant natural language processing tasks.
Understanding and improving transformer from a multi-particle dynamic system point of view
Yiping Lu, Zhuohan Li, Di He, Zhiqing Sun, Bin Dong, Tao Qin, Liwei Wang, and Tie-Yan Liu. 2019 · 1906
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Robust learning with jacobian regularization
Judy Hoffman, Daniel A Roberts, and Sho Yaida. 2019 · 1908
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
Scheduled drophead: A regularization method for transformer models
Wangchunshu Zhou, Tao Ge, Furu Wei, Ming Zhou, and Ke Xu. 2020 · 1980
Earlier work this paper cites.
A simple weight decay can improve generalization
Anders Krogh and John A Hertz. 1992 · 1992
Earlier work this paper cites.
The TREC-8 question answering track evaluation
Ellen M. Voorhees and Dawn M. Tice. 1999 · 1999
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
The implicit and explicit regularization effects of dropout
Colin Wei, Sham Kakade, and Tengyu Ma. 2020 · 2002
Earlier work this paper cites.
Multi-branch attentive transformer
Yang Fan, Shufang Xie, Yingce Xia, Lijun Wu, Tao Qin, Xiang-Yang Li, and Tie-Yan Liu. 2020b · 2006
Earlier work this paper cites.
Alleviating the inequality of attention heads for neural machine translation
Zewei Sun, Shujian Huang, Xinyu Dai, and Jiajun Chen. 2020 · 2009
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E Hinton. 2010 · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Earlier work this paper cites.
Deep unordered composition rivals syntactic methods for text classification
Mohit Iyyer, Varun Manjunatha, Jordan Boyd-Graber, and Hal Daumé III. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Jake Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016 · 2016
Cited alongside, same era.
A theoretically grounded application of dropout in recurrent neural networks
Yarin Gal and Zoubin Ghahramani. 2016 · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Deep networks with stochastic depth
Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q. Weinberger. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016b · 2016
Cited alongside, same era.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019 · 2019
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Soft contextual data augmentation for neural machine translation
Fei Gao, Jinhua Zhu, Lijun Wu, Yingce Xia, Tao Qin, Xueqi Cheng, Wengang Zhou, and Tie-Yan Liu. 2019 · 2019
Later among the works it cites.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Later among the works it cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Karim Ahmed, Nitish Shirish Keskar, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Supervised learning of universal sentence representations from natural language inference data
Alexis Conneau, Douwe Kiela, Holger Schwenk, Loïc Barrault, and Antoine Bordes. 2017 · 2017
Cited alongside, same era.
Deep pyramid convolutional neural networks for text categorization
Rie Johnson and Tong Zhang. 2017 · 2017
Cited alongside, same era.
Zoneout: Regularizing rnns by randomly preserving hidden activations
David Krueger, Tegan Maharaj, János Kramár, Mohammad Pezeshki, Nicolas Ballas, Nan Rosemary Ke, Anirudh Goyal, Yoshua Bengio, Aaron C. Courville, and Christopher J. Pal. 2017 · 2017
Cited alongside, same era.
Fractalnet: Ultra-deep neural networks without residuals
Gustav Larsson, Michael Maire, and Gregory Shakhnarovich. 2017 · 2017
Cited alongside, same era.
Regularizing neural networks by penalizing confident output distributions
Gabriel Pereyra, George Tucker, Jan Chorowski, Lukasz Kaiser, and Geoffrey E. Hinton. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019a · 2019
Later among the works it cites.
Improving neural language modeling via adversarial training
Dilin Wang, ChengYue Gong, and Qiang Liu. 2019b · 2019
Later among the works it cites.
EDA: easy data augmentation techniques for boosting performance on text classification tasks
Jason W. Wei and Kai Zou. 2019 · 2019
Later among the works it cites.
Pay less attention with lightweight and dynamic convolutions
Felix Wu, Angela Fan, Alexei Baevski, Yann Dauphin, and Michael Auli. 2019 · 2019
Later among the works it cites.
Tied transformers: Neural machine translation with shared encoder and decoder
Yingce Xia, Tianyu He, Xu Tan, Fei Tian, Di He, and Tao Qin. 2019 · 2019
Later among the works it cites.
Reducing transformer depth on demand with structured dropout
Angela Fan, Edouard Grave, and Armand Joulin. 2020a · 2020
Later among the works it cites.
Guided transformer: Leveraging multiple external sources for representation learning in conversational search
Helia Hashemi, Hamed Zamani, and W Bruce Croft. 2020 · 2020
Later among the works it cites.
Reformer: The efficient transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. 2020 · 2020
Later among the works it cites.
Sequence generation with mixed representations
Lijun Wu, Shufang Xie, Yingce Xia, Yang Fan, Tao Qin, Jianhuang Lai, and Tie-Yan Liu. 2020 · 2020
Later among the works it cites.
Unsupervised data augmentation for consistency training
Qizhe Xie, Zihang Dai, Eduard Hovy, Thang Luong, and Quoc Le. 2020 · 2020
Later among the works it cites.
Incorporating BERT into neural machine translation
Jinhua Zhu, Yingce Xia, Lijun Wu, Di He, Tao Qin, Wengang Zhou, Houqiang Li, and Tie-Yan Liu. 2020 · 2020
Later among the works it cites.
{IOT}: Instance-wise layer reordering for transformer structures
Jinhua Zhu, Lijun Wu, Yingce Xia, Shufang Xie, Tao Qin, Wengang Zhou, Houqiang Li, and Tie-Yan Liu. 2021 · 2021
Closest in time.