A decomposable attention model for natural language inference
Ankur P. Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit. 2016 · 2016
Later among the works it cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Original
Tim Salimans and Diederik P. Kingma. 2016 · 2016
Later among the works it cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Original
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016 · 2016
Later among the works it cites.
Deep recurrent models with fast-forward connections for neural machine translation
Original
Jie Zhou, Ying Cao, Xuguang Wang, Peng Li, and Wei Xu. 2016 · 2016
Later among the works it cites.
Massive exploration of neural machine translation architectures
Denny Britz, Anna Goldie, Minh-Thang Luong, and Quoc Le. 2017 · 2017
Later among the works it cites.
Checkpoint ensembles: Ensemble methods from a single training process
Original
Hugh Chen, Scott Lundberg, and Su-In Lee. 2017 · 2017
Later among the works it cites.
Stronger baselines for trustable results in neural machine translation
Michael Denkowski and Graham Neubig. 2017 · 2017
Later among the works it cites.
Sharp models on dull hardware: Fast and accurate neural machine translation decoding on the cpu
Jacob Devlin. 2017 · 2017
Later among the works it cites.
Convolutional sequence to sequence learning
Original
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin. 2017 · 2017
Later among the works it cites.
Accurate, large minibatch SGD: training imagenet in 1 hour
Original
Priya Goyal, Piotr Dollár, Ross B. Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. 2017 · 2017
Later among the works it cites.
In-datacenter performance analysis of a tensor processing unit
Original
Norman P. Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, and et al. 2017 · 2017
Later among the works it cites.
Attention is all you need
Original
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.