Fetching the paper…
Reading the bibliography…
Deep neural networks (DNNs) continue to make significant advances, solving tasks from image classification to translation or reinforcement learning.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Europarl: A Parallel Corpus for Statistical Machine Translation
Philipp Koehn · 2005
Earlier work this paper cites.
Ordrec: an ordinal model for predicting personalized item rating distributions
Yehuda Koren and Joe Sill · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Do deep nets really need to be deep?
Lei Jimmy Ba and Rich Caruana · 2013
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron C. Courville · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann Dauphin, Razvan Pascanu, Çaglar Gülçehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Earlier work this paper cites.
Matrix recovery from quantized and corrupted measurements
Andrew S Lan, Christoph Studer, and Richard G Baraniuk · 2014
Earlier work this paper cites.
Binaryconnect: Training deep neural networks with binary weights during propagations
Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David · 2015
Earlier work this paper cites.
Song Han, Huizi Mao, and William J. Dally · 2015
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D. Manning · 2015
Earlier work this paper cites.
Learning using privileged information: similarity control and knowledge transfer
Vladimir Vapnik and Rauf Izmailov · 2015
Cited alongside, same era.
Qsgd: Randomized quantization for communication-optimal stochastic gradient descent
Dan Alistarh, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2016
Cited alongside, same era.
Noisy activation functions
Caglar Gulcehre, Marcin Moczulski, Misha Denil, and Yoshua Bengio · 2016
Cited alongside, same era.
Hardware-oriented approximation of convolutional neural networks
Philipp Gysel, Mohammad Motamedi, and Soheil Ghiasi · 2016
Cited alongside, same era.
Quantized neural networks: Training neural networks with low precision weights and activations
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio · 2016
Simple and efficient learning using privileged information
Xinxing Xu, Joey Tianyi Zhou, IvorW Tsang, Zheng Qin, Rick Siow Mong Goh, and Yong Liu · 2016
Later among the works it cites.
Wide Residual Networks
S. Zagoruyko and N. Komodakis · 2016
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Later among the works it cites.
Chenzhuo Zhu, Song Han, Huizi Mao, and William J. Dally · 2016
Later among the works it cites.
https://github.com/OpenNMT/OpenNMT-py
Opennmt integration testing · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Squeezenet: Alexnet-level accuracy with 50x fewer parameters and¡ 0.5 mb model size
Forrest N Iandola, Song Han, Matthew W Moskewicz, Khalid Ashraf, William J Dally, and Kurt Keutzer · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Recurrent neural networks with limited numerical precision
Joachim Ott, Zhouhan Lin, Ying Zhang, Shih-Chii Liu, and Yoshua Bengio · 2016
Cited alongside, same era.
Xnor-net: Imagenet classification using binary convolutional neural networks
Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Do Deep Convolutional Nets Really Need to be Deep and Convolutional?
G. Urban, K. J. Geras, S. Ebrahimi Kahou, O. Aslan, S. Wang, R. Caruana, A. Mohamed, M. Philipose, and M. Richardson · 2016
Cited alongside, same era.
http://www.statmt.org/moses/?n=moses.baseline
Moses baseline · 2017
Later among the works it cites.
OpenNMT: Open-Source Toolkit for Neural Machine Translation
G. Klein, Y. Kim, Y. Deng, J. Senellart, and A. M. Rush · 2017
Later among the works it cites.
Training quantized nets: A deeper understanding
Hao Li, Soham De, Zheng Xu, Christoph Studer, Hanan Samet, and Tom Goldstein · 2017
Later among the works it cites.
Ternary neural networks with fine-grained quantization
Naveen Mellempudi, Abhisek Kundu, Dheevatsa Mudigere, Dipankar Das, Bharat Kaul, and Pradeep Dubey · 2017
Later among the works it cites.
WRPN: wide reduced-precision networks
Asit K. Mishra, Eriko Nurvitadhi, Jeffrey J. Cook, and Debbie Marr · 2017
Later among the works it cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
Zipml: Training linear models with end-to-end low precision, and a little bit of deep learning
Hantian Zhang, Jerry Li, Kaan Kara, Dan Alistarh, Ji Liu, and Ce Zhang · 2017
Later among the works it cites.