Fetching the paper…
Reading the bibliography…
The Transformer is widely used in natural language processing tasks.
A mean field theory of batch normalization
Yang, G., Pennington, J., Rao, V., Sohl-Dickstein, J., and Schoenholz, S. S · 1902
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R., and Le, Q. V · 1906
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 1907
Earlier work this paper cites.
On the variance of the adaptive learning rate and beyond
Liu, L., Jiang, H., He, P., Chen, W., Liu, X., Gao, J., and Han, J · 1908
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, W. B. and Brockett, C · 2005
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Koehn, P., Hoang, H., Birch, A., Callison-Burch, C., Federico, M., Bertoldi, N., Cowan, B., Shen, W., Moran, C., Zens, R., et al · 2007
Earlier work this paper cites.
The fifth PASCAL recognizing textual entailment challenge
Bentivogli, L., Dagan, I., Dang, H. T., Giampiccolo, D., and Magnini, B · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop, coursera: Neural networks for machine learning
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Zeiler, M. D · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2015
Earlier work this paper cites.
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Lei Ba, J., Kiros, J. R., and Hinton, G. E · 2016
Cited alongside, same era.
An overview of gradient descent optimization algorithms
Ruder, S · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2017
Cited alongside, same era.
Language modeling with gated convolutional networks
Opennmt: Neural machine translation toolkit
Klein, G., Kim, Y., Deng, Y., Nguyen, V., Senellart, J., and Rush, A · 2018
Later among the works it cites.
Training tips for the transformer model
Popel, M. and Bojar, O · 2018
Later among the works it cites.
Tensor2tensor for neural machine translation
Vaswani, A., Bengio, S., Brevdo, E., Chollet, F., Gomez, A. N., Gouws, S., Jones, L., Kaiser, L., Kalchbrenner, N., Parmar, N., Sepassi, R., Shazeer, N., and Uszkoreit, J · 2018
Later among the works it cites.
Dynamical isometry and a mean field theory of cnns: How to train 10,000-layer vanilla convolutional neural networks
Xiao, L., Bahri, Y., Sohl-Dickstein, J., Schoenholz, S., and Pennington, J · 2018
Later among the works it cites.
Imagenet training in minutes
You, Y., Zhang, Z., Hsieh, C.-J., Demmel, J., and Keutzer, K · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dauphin, Y. N., Fan, A., Auli, M., and Grangier, D · 2017
Cited alongside, same era.
Convolutional sequence to sequence learning
Gehring, J., Auli, M., Grangier, D., Yarats, D., and Dauphin, Y. N · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Cited alongside, same era.
Mask r-cnn
He, K., Gkioxari, G., Dollár, P., and Girshick, R · 2017
Cited alongside, same era.
Deep neural networks as gaussian processes
Lee, J., Bahri, Y., Novak, R., Schoenholz, S. S., Pennington, J., and Sohl-Dickstein, J · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Character-level language modeling with deeper self-attention
Al-Rfou, R., Choe, D., Constant, N., Guo, M., and Jones, L · 2018
Cited alongside, same era.
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Later among the works it cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Cohen, W. W., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2019
Later among the works it cites.
Bag of tricks for image classification with convolutional neural networks
He, T., Zhang, Z., Zhang, H., Zhang, Z., Xie, J., and Li, M · 2019
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J., Xiao, L., Schoenholz, S. S., Bahri, Y., Sohl-Dickstein, J., and Pennington, J · 2019
Later among the works it cites.
Understanding and improving transformer from a multi-particle dynamic system point of view
Lu, Y., Li, Z., He, D., Sun, Z., Dong, B., Qin, T., Wang, L., and Liu, T.-Y · 2019
Later among the works it cites.
Transformers without tears: Improving the normalization of self-attention
Nguyen, T. Q. and Salazar, J · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Later among the works it cites.
High-dimensional statistics: A non-asymptotic viewpoint , volume 48
Wainwright, M. J · 2019
Later among the works it cites.
Learning deep transformer models for machine translation
Wang, Q., Li, B., Xiao, T., Zhu, J., Li, C., Wong, D. F., and Chao, L. S · 2019
Later among the works it cites.
Yang, G · 2019
Later among the works it cites.
Fixup initialization: Residual learning without normalization
Zhang, H., Dauphin, Y. N., and Ma, T · 2019
Later among the works it cites.