Fetching the paper…
Reading the bibliography…
The performance of a deep neural network is highly dependent on its training, and finding better local optimal solutions is the goal of many optimization algorithms.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Learning internal representations by error propagation
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1985
Earlier work this paper cites.
The effects of adding noise during backpropagation training on a generalization performance
G. An · 1996
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, P. Haffner, et al · 1998
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
N. Qian · 1999
Earlier work this paper cites.
Stochastic neighbor embedding
G. E. Hinton and S. T. Roweis · 2003
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
M. Zinkevich · 2003
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky, G. Hinton, et al · 2009
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman · 2010
Earlier work this paper cites.
Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion
P. Vincent, H. Larochelle, I. Lajoie, Y. Bengio, and P.-A. Manzagol · 2010
Earlier work this paper cites.
Practical variational inference for neural networks
A. Graves · 2011
Earlier work this paper cites.
Semantic contours from inverse detectors
B. Hariharan, P. Arbeláez, L. Bourdev, S. Maji, and J. Malik · 2011
Earlier work this paper cites.
Efficiency of coordinate descent methods on huge-scale optimization problems
Y. Nesterov · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
M. D. Zeiler · 2012
Cited alongside, same era.
Random coordinate descent algorithms for multi-agent convex optimization over networks
I. Necoara · 2013
Cited alongside, same era.
Regularization of neural networks using dropconnect
L. Wan, M. Zeiler, S. Zhang, Y. Le Cun, and R. Fergus · 2013
Cited alongside, same era.
Fast dropout training
S. Wang and C. Manning · 2013
Cited alongside, same era.
Semantic image segmentation with deep convolutional nets and fully connected crfs
L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille · 2014
Cited alongside, same era.
Efficient piecewise training of deep structured models for semantic segmentation
G. Lin, C. Shen, A. Van Den Hengel, and I. Reid · 2016
Later among the works it cites.
Ssd: Single shot multibox detector
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg · 2016
Later among the works it cites.
Disturblabel: Regularizing cnn on the loss layer
L. Xie, J. Wang, Z. Wei, M. Wang, and Q. Tian · 2016
Later among the works it cites.
Noisy softmax: Improving the generalization ability of dcnn via postponing the early softmax saturation
B. Chen, W. Deng, and J. Du · 2017
Later among the works it cites.
Densely connected convolutional networks
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
A. C. Wilson, R. Roelofs, M. Stern, N. Srebro, and B. Recht · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Cited alongside, same era.
Weight uncertainty in neural networks
C. Blundell, J. Cornebise, K. Kavukcuoglu, and D. Wierstra · 2015
Cited alongside, same era.
Fully convolutional networks for semantic segmentation
J. Long, E. Shelhamer, and T. Darrell · 2015
Cited alongside, same era.
Adding gradient noise improves learning for very deep networks
A. Neelakantan, L. Vilnis, Q. V. Le, I. Sutskever, L. Kaiser, K. Kurach, and J. Martens · 2015
Cited alongside, same era.
Coordinate descent algorithms
S. J. Wright · 2015
Cited alongside, same era.
Noisy activation functions
C. Gulcehre, M. Moczulski, M. Denil, and Y. Bengio · 2016
Cited alongside, same era.
Later among the works it cites.
Pyramid scene parsing network
H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia · 2017
Later among the works it cites.
Introducing noise in decentralized training of neural networks
L. Adilova, N. Paul, and P. Schlicht · 2018
Later among the works it cites.
A pid controller approach for stochastic optimization of deep networks
W. An, H. Wang, Q. Sun, J. Xu, Q. Dai, and L. Zhang · 2018
Later among the works it cites.
Encoder-decoder with atrous separable convolution for semantic image segmentation
L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, and H. Adam · 2018
Later among the works it cites.
On the convergence of adam and beyond
S. J. Reddi, S. Kale, and S. Kumar · 2018
Later among the works it cites.
On the variance of the adaptive learning rate and beyond
L. Liu, H. Jiang, P. He, W. Chen, X. Liu, J. Gao, and J. Han · 2019
Closest in time.
Adaptive gradient methods with dynamic bound of learning rate
L. Luo, Y. Xiong, Y. Liu, and X. Sun · 2019
Closest in time.