Fetching the paper…
Reading the bibliography…
As traditional neural network consumes a significant amount of computing resources during back propagation, \citet{Sun2017mePropSB} propose a simple yet effective technique to alleviate this problem.
Supersab: fast adaptive back propagation with good scaling properties
Tollenaere, Tom · 1990
Earlier work this paper cites.
A direct adaptive method for faster backpropagation learning: The rprop algorithm
Riedmiller, Martin and Braun, Heinrich · 1993
Earlier work this paper cites.
Natural image statistics and efficient coding
Olshausen, Bruno A and Field, David J · 1996
Earlier work this paper cites.
Efficient learning of sparse representations with an energy-based model
Poultney, Christopher, Chopra, Sumit, Cun, Yann L, et al · 2007
Earlier work this paper cites.
Modeling latent-dynamic in shallow parsing: A latent conditional model with improved inference
Sun, Xu, Morency, Louis-Philippe, Okanohara, Daisuke, and Tsujii, Jun’ichi · 2008
Earlier work this paper cites.
A large scale ranker-based system for search query spelling correction
Gao, Jianfeng, Li, Xiaolong, Micol, Daniel, Quirk, Chris, and Sun, Xu · 2010
Earlier work this paper cites.
Mnist handwritten digit database
LeCun, Yann, Cortes, Corinna, and Burges, Christopher JC · 2010
Earlier work this paper cites.
Learning phrase-based spelling error models from clickthrough data
Sun, Xu, Gao, Jianfeng, Micol, Daniel, and Quirk, Chris · 2010
Earlier work this paper cites.
Deep sparse rectifier neural networks
Glorot, Xavier, Bordes, Antoine, and Bengio, Yoshua · 2011
Cited alongside, same era.
A convergence analysis of log-linear training
Wiesler, Simon and Ney, Hermann · 2011
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E · 2012
Cited alongside, same era.
Efficient backprop
LeCun, Yann A, Bottou, Léon, Orr, Genevieve B, and Müller, Klaus-Robert · 2012
Cited alongside, same era.
Fast online training with frequency-adaptive learning rates for chinese word segmentation and new word detection
Sun, Xu, Wang, Houfeng, and Li, Wenjie · 2012
Cited alongside, same era.
On using very large target vocabulary for neural machine translation
Jean, Sébastien, Cho, Kyunghyun, Memisevic, Roland, and Bengio, Yoshua · 2014
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, Nitish, Hinton, Geoffrey E, Krizhevsky, Alex, Sutskever, Ilya, and Salakhutdinov, Ruslan · 2014
Later among the works it cites.
Deepface: Closing the gap to human-level performance in face verification
Taigman, Yaniv, Yang, Ming, Ranzato, Marc’Aurelio, and Wolf, Lior · 2014
Later among the works it cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, Martín, Agarwal, Ashish, Barham, Paul, Brevdo, Eugene, Chen, Zhifeng, Citro, Craig, Corrado, Greg S., Davis, Andy, Dean, Jeffrey, Devin, Matthieu, Ghemawat, Sanjay, Goodfellow, Ian, Harp, Andrew, Irving, Geoffrey, Isard, Michael, Jia, Yangqing, Jozefowicz, Rafal, Kaiser, Lukasz, Kudlur, Manjunath, Levenberg, Josh, Mané, Dan, Monga, Rajat, Moore, Sherry, Murray, Derek, Olah, Chris, Schuster, Mike, Shlens, Jonathon, Steiner, Benoit, Sutskever, Ilya, Talwar, Kunal, Tucker, Paul, Vanhoucke, Vincent, Vasudevan, Vijay, Viégas, Fernanda, Vinyals, Oriol, Warden, Pete, Wattenberg, Martin, Wicke, Martin, Yu, Yuan, and Zheng, Xiaoqiang · 2015
Later among the works it cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, Sergey and Szegedy, Christian · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2014
Cited alongside, same era.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Seide, Frank, Fu, Hao, Droppo, Jasha, Li, Gang, and Yu, Dong · 2014
Cited alongside, same era.
Dryden, Nikoli, Jacobs, Sam Ade, Moon, Tim, and Van Essen, Brian · 2016
Later among the works it cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, Noam, Mirhoseini, Azalia, Maziarz, Krzysztof, Davis, Andy, Le, Quoc, Hinton, Geoffrey, and Dean, Jeff · 2017
Closest in time.
meprop: Sparsified back propagation for accelerated deep learning with reduced overfitting
Sun, Xu, Ren, Xuancheng, Ma, Shuming, and Wang, Houfeng · 2017
Closest in time.