Google’s neural machine translation system: Bridging the gap between human and machine translation
Original
Wu, Yonghui, Schuster, Mike, Chen, Zhifeng, Le, Quoc V., Norouzi, Mohammad, Macherey, Wolfgang, Krikun, Maxim, Cao, Yuan, Gao, Qin, Macherey, Klaus, Klingner, Jeff, Shah, Apurva, Johnson, Melvin, Liu, Xiaobing, Kaiser, Lukasz, Gouws, Stephan, Kato, Yoshikiyo, Kudo, Taku, Kazawa, Hideto, Stevens, Keith, Kurian, George, Patil, Nishant, Wang, Wei, Young, Cliff, Smith, Jason, Riesa, Jason, Rudnick, Alex, Vinyals, Oriol, Corrado, Greg, Hughes, Macduff, and Dean, Jeffrey · 2016
Later among the works it cites.
Wide residual networks
Original
Zagoruyko, Sergey and Komodakis, Nikos · 2016
Later among the works it cites.
Distributed second-order optimization using Kronecker-factored approximations
Ba, Jimmy, Grosse, Roger, and Martens, James · 2017
Closest in time.
Google vizier: A service for black-box optimization
Golovin, Daniel, Solnik, Benjamin, Moitra, Subhodeep, Kochanski, Greg, Karro, John Elliot, and Sculley, D · 2017
Closest in time.
Learning to optimize neural nets
Original
Li, Ke and Malik, Jitendra · 2017
Closest in time.
SGDR: stochastic gradient descent with restarts
Loshchilov, Ilya and Hutter, Frank · 2017
Closest in time.
On the state of the art of evaluation in neural language models
Original
Melis, Gabor, Dyer, Chris, and Blunsom, Phil · 2017
Closest in time.
Using the output embedding to improve language models
Original
Press, Ofir and Wolf, Lior · 2017
Closest in time.
Optimization as a model for few-shot learning
Ravi, Sachin and Larochelle, Hugo · 2017
Closest in time.
Proximal policy optimization algorithms
Schulman, John, Wolski, Filip, Dhariwal, Prafulla, Radford, Alec, and Klimov, Oleg · 2017
Closest in time.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, Noam, Mirhoseini, Azalia, Maziarz, Krzysztof, Davis, Andy, Le, Quoc, Hinton, Geoffrey, and Dean, Jeff · 2017
Closest in time.
Learned optimizers that scale and generalize
Wichrowska, Olga, Maheswaranathan, Niru, Hoffman, Matthew W., Colmenarejo, Sergio Gomez, Denil, Misha, de Freitas, Nando, and Sohl-Dickstein, Jascha · 2017
Closest in time.
Understanding deep learning requires rethinking generalization
Zhang, Chiyuan, Bengio, Samy, Hardt, Moritz, Recht, Benjamin, and Vinyals, Oriol · 2017
Closest in time.
Neural Architecture Search with reinforcement learning
Zoph, Barret and Le, Quoc V · 2017
Closest in time.
Learning transferable architectures for scalable image recognition
Original
Zoph, Barret, Vasudevan, Vijay, Shlens, Jonathon, and Le, Quoc V · 2017
Closest in time.