Fetching the paper…
Reading the bibliography…
Optimal parameter initialization remains a crucial problem for neural network training.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Examples of adaptive mcmc
Gareth O Roberts and Jeffrey S Rosenthal · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples (2014)
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2014
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Alex Krizhevsky · 2014
Earlier work this paper cites.
Recurrent neural network regularization
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Dmytro Mishkin and Jiri Matas · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Cited alongside, same era.
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
Enhanced lstm for natural language inference
Qian Chen, Xiaodan Zhu, Zhen-Hua Ling, Si Wei, Hui Jiang, and Diana Inkpen · 2017
Cited alongside, same era.
Charles H Martin and Michael W Mahoney · 2017
Later among the works it cites.
Cyclical learning rates for training neural networks
Leslie N Smith · 2017
Later among the works it cites.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, and Quoc V Le · 2017
Later among the works it cites.
Second-order optimization for non-convex machine learning: An empirical study
Peng Xu, Farbod Roosta-Khorasan, and Michael W Mahoney · 2017
Later among the works it cites.
Integrated model, batch and domain parallelism in training neural networks
Amir Gholami, Ariful Azad, Peter Jin, Kurt Keutzer, and Aydin Buluc · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aditya Devarakonda, Maxim Naumov, and Michael Garland · 2017
Cited alongside, same era.
On an adaptive preconditioned crank–nicolson mcmc algorithm for infinite dimensional bayesian inference
Zixi Hu, Zhewei Yao, and Jinglai Li · 2017
Cited alongside, same era.
Snapshot ensembles: Train 1, get m for free
Gao Huang, Yixuan Li, Geoff Pleiss, Zhuang Liu, John E Hopcroft, and Kilian Q Weinberger · 2017
Cited alongside, same era.
Three factors influencing minima in sgd
Stanis l · 2017
Cited alongside, same era.
Large batch size training of neural networks with adversarial training and second-order information
Zhewei Yao, Amir Gholami, Kurt Keutzer, and Michael Mahoney
Cited in the paper.
Hessian-based analysis of large batch training and robustness to adversaries
Zhewei Yao, Amir Gholami, Qi Lei, Kurt Keutzer, and Michael W Mahoney
Cited in the paper.
Closest in time.
Charles H Martin and Michael W Mahoney · 2018
Closest in time.
A Bayesian perspective on generalization and Stochastic Gradient Descent
Samuel L Smith and Quoc V Le · 2018
Closest in time.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Closest in time.