Fetching the paper…
Reading the bibliography…
We describe a novel family of models of multi- layer feedforward neural networks in which the activation functions are encoded via penalties in the training problem.
Combining the predictions of multiple classifiers: Using competitive learning to initialize neural networks
Maclin, Richard and Shavlik, Jude W · 1995
Earlier work this paper cites.
Effiicient backprop
LeCun, Yann, Bottou, Léon, Orr, Genevieve B., and Müller, Klaus-Robert · 1998
Earlier work this paper cites.
Fast Newton-type methods for the least squares nonnegative matrix approximation problem
Kim, Dongmin, Sra, Suvrit, and Dhillon, Inderjit S · 2007
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, Xavier and Bengio, Yoshua · 2010
Earlier work this paper cites.
MNIST handwritten digit database
LeCun, Yann and Cortes, Corinna · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, John C., Hazan, Elad, and Singer, Yoram · 2011
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, Ilya, Martens, James, Dahl, George, and Hinton, Geoffrey · 2013
Earlier work this paper cites.
Distributed optimization of deeply nested systems
Carreira-Perpinan, Miguel and Wang, Weiran · 2014
Cited alongside, same era.
Algorithms for nonnegative matrix and tensor factorizations: A unified view based on block coordinate descent framework
Kim, Jingu, He, Yunlong, and Park, Haesun · 2014
Cited alongside, same era.
Sketching as a tool for numerical linear algebra
Woodruff, David P et al · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, Martín, Agarwal, Ashish, Barham, Paul, Brevdo, Eugene, Chen, Zhifeng, Citro, Craig, Corrado, Greg S., Davis, Andy, Dean, Jeffrey, Devin, Matthieu, Ghemawat, Sanjay, Goodfellow, Ian, Harp, Andrew, Irving, Geoffrey, Isard, Michael, Jia, Yangqing, Jozefowicz, Rafal, Kaiser, Lukasz, Kudlur, Manjunath, Levenberg, Josh, Mané, Dan, Monga, Rajat, Moore, Sherry, Murray, Derek, Olah, Chris, Schuster, Mike, Shlens, Jonathon, Steiner, Benoit, Sutskever, Ilya, Talwar, Kunal, Tucker, Paul, Vanhoucke, Vincent, Vasudevan, Vijay, Viégas, Fernanda, Vinyals, Oriol, Warden, Pete, Wattenberg, Martin, Wicke, Martin, Yu, Yuan, and Zheng, Xiaoqiang · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik P. and Ba, Jimmy · 2015
Later among the works it cites.
Trusting svm for piecewise linear cnns
Berrada, Leonard, Zisserman, Andrew, and Kumar, M. Pawan · 2016
Later among the works it cites.
Iterative Hessian sketch: Fast and accurate solution approximation for constrained least-squares
Pilanci, Mert and Wainwright, Martin J · 2016
Later among the works it cites.
Training neural networks without gradients: A scalable admm approach
Taylor, Gavin, Burmeister, Ryan, Xu, Zheng, Singh, Bharat, Patel, Ankit, and Goldstein, Tom · 2016
Later among the works it cites.
PCA-initialized deep neural networks applied to document image analysis
Seuret, Mathias, Alberti, Michele, Ingold, Rolf, and Liwicki, Marcus · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bubeck, Sébastien · 2015
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, Kaiming, Zhang, Xiangyu, Ren, Shaoqing, and Sun, Jian · 2015
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
Wilson, Ashia C., Roelofs, Rebecca, Stern, Mitchell, Srebro, Nati, and Recht, Benjamin · 2017
Later among the works it cites.