Character-aware neural language models
Yoon Kim, Yacine Jernite, David Sontag, and Alexander M Rush · 2016
Cited alongside, same era.
Large-margin softmax loss for convolutional neural networks
Weiyang Liu, Yandong Wen, Zhiding Yu, and Meng Yang · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Original
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Cited alongside, same era.
Massive exploration of neural machine translation architectures
Denny Britz, Anna Goldie, Minh-Thang Luong, and Quoc Le · 2017
Cited alongside, same era.
Language modeling with gated convolutional networks
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier · 2017
Cited alongside, same era.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N Dauphin · 2017
Cited alongside, same era.
Improving neural language models with a continuous cache
Edouard Grave, Armand Joulin, and Nicolas Usunier · 2017
Cited alongside, same era.
Tying word vectors and word classifiers: A loss framework for language modeling
Hakan Inan, Khashayar Khosravi, and Richard Socher · 2017
Cited alongside, same era.
Stochastic gradient descent as approximate bayesian inference
Stephan Mandt, Matthew D Hoffman, and David M Blei · 2017
Cited alongside, same era.
Dual inference for machine learning
Yingce Xia, Jiang Bian, Tao Qin, Nenghai Yu, and Tie-Yan Liu
Cited in the paper.
Dual supervised learning
Yingce Xia, Tao Qin, Wei Chen, Jiang Bian, Nenghai Yu, and Tie-Yan Liu
Cited in the paper.