Fetching the paper…
Reading the bibliography…
We present new algorithms for adaptively learning artificial neural networks.
Optimal brain damage
LeCun, Yann, Denker, John S., and Solla, Sara A · 1990
Earlier work this paper cites.
On the convergence of coordinate descent method for convex differentiable minimization
Luo, Zhi-Quan and Tseng, Paul · 1992
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Freund, Yoav and Schapire, Robert E · 1997
Earlier work this paper cites.
A structural learning algorithm for multi-layered neural networks
Kotani, Manabu, Kajiki, Akihiro, and Akazawa, Kenzo · 1997
Earlier work this paper cites.
Constructive algorithms for structure learning in feedforward neural networks for regression problems
Kwok, Tin-Yau and Yeung, Dit-Yan · 1997
Earlier work this paper cites.
The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network
Bartlett, Peter L · 1998
Earlier work this paper cites.
Modelling with constructive backpropagation
Lehtokangas, Mikko · 1999
Earlier work this paper cites.
On the convergence of leveraging
Rätsch, Gunnar, Mika, Sebastian, and Warmuth, Manfred K · 2001
Earlier work this paper cites.
Rademacher and Gaussian complexities: Risk bounds and structural results
Bartlett, Peter L. and Mendelson, Shahar · 2002
Earlier work this paper cites.
Empirical margin distributions and bounding the generalization error of combined classifiers
Koltchinskii, Vladmir and Panchenko, Dmitry · 2002
Earlier work this paper cites.
A constructive algorithm for training cooperative neural network ensembles
Islam, Md Monirul, Yao, Xin, and Murase, Kazuyuki · 2003
Earlier work this paper cites.
Tuning of the structure and parameters of a neural network using an improved genetic algorithm
Leung, Frank HF, Lam, Hak-Keung, Ling, Sai-Ho, and Tam, Peter KS · 2003
Earlier work this paper cites.
A new strategy for adaptively constructing multilayer feedforward neural networks
Ma, Liying and Khorasani, Khashayar · 2003
Earlier work this paper cites.
An integrated growing-pruning method for feedforward network training
Narasimha, Pramod L, Delashmit, Walter H, Manry, Michael T, Li, Jiang, and Maldonado, Francisco · 2008
Earlier work this paper cites.
A new adaptive merging and growing algorithm for designing artificial neural networks
Islam, MohamOBmad, Sattar, Abdul, Amin, Farnaz, Yao, Xin, and Murase, Kazuyuki · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, Alex · 2009
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
Bergstra, James S, Bardenet, Rémi, Bengio, Yoshua, and Kégl, Balázs · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E · 2012
Cited alongside, same era.
Practical Bayesian Optimization of Machine Learning Algorithms
Snoek, Jasper, Larochelle, Hugo, and Adams, Ryan P · 2012
Cited alongside, same era.
A structure optimisation algorithm for feedforward neural network construction
Han, Hong-Gui and Qiao, Jun-Fei · 2013
Cited alongside, same era.
On the difficulty of training recurrent neural networks
Pascanu, Razvan, Mikolov, Tomas, and Bengio, Yoshua · 2013
Cited alongside, same era.
Provable bounds for learning some deep representations
Arora, Sanjeev, Bhaskara, Aditya, Ge, Rong, and Ma, Tengyu · 2014
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, Moritz, Recht, Benjamin, and Singer, Yoram · 2015
Later among the works it cites.
Deep residual learning for image recognition
He, Kaiming, Zhang, Xiangyu, Ren, Shaoqing, and Sun, Jian · 2015
Later among the works it cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, Sergey and Szegedy, Christian · 2015
Later among the works it cites.
Generalization bounds for neural networks through tensor factorization
Janzamin, Majid, Sedghi, Hanie, and Anandkumar, Anima · 2015
Later among the works it cites.
Asynchronous parallel stochastic gradient for nonconvex optimization
Lian, Xiangru, Huang, Yijun, Li, Yuncheng, and Liu, Ji · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Choromanska, Anna, Henaff, Mikael, Mathieu, Michael, Arous, Gérard Ben, and LeCun, Yann · 2014
Cited alongside, same era.
Deep boosting
Cortes, Corinna, Mohri, Mehryar, and Syed, Umar · 2014
Cited alongside, same era.
Multi-class deep boosting
Kuznetsov, Vitaly, Mohri, Mehryar, and Syed, Umar · 2014
Cited alongside, same era.
On the computational efficiency of training neural networks
Livni, Roi, Shalev-Shwartz, Shai, and Shamir, Ohad · 2014
Cited alongside, same era.
Explorations on high dimensional landscapes
Sagun, Levent, Guney, V Ugur, Arous, Gerard Ben, and LeCun, Yann · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Sutskever, Ilya, Vinyals, Oriol, and Le, Quoc V · 2014
Cited alongside, same era.
Neural architecture search with reinforcement learning
Zoph, Barret and Le, Quoc V · 2014
Cited alongside, same era.
Later among the works it cites.
Norm-based capacity control in neural networks
Neyshabur, Behnam, Tomioka, Ryota, and Srebro, Nathan · 2015
Later among the works it cites.
Going deeper with convolutions
Szegedy, Christian, Liu, Wei, Jia, Yangqing, Sermanet, Pierre, Reed, Scott E., Anguelov, Dragomir, Erhan, Dumitru, Vanhoucke, Vincent, and Rabinovich, Andrew · 2015
Later among the works it cites.
ℓ _ 1 \ell\_1 -regularized neural networks are improperly learnable in polynomial time
Zhang, Yuchen, Lee, Jason D, and Jordan, Michael I · 2015
Later among the works it cites.
Learning the number of neurons in deep networks
Alvarez, Jose M and Salzmann, Mathieu · 2016
Closest in time.
Designing neural network architectures using reinforcement learning
Baker, Bowen, Gupta, Otkrist, Naik, Nikhil, and Raskar, Ramesh · 2016
Closest in time.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Daniely, Amit, Frostig, Roy, and Singer, Yoram · 2016
Closest in time.
Hypernetworks
Ha, David, Dai, Andrew M., and Le, Quoc V · 2016
Closest in time.
Densely connected convolutional networks
Huang, Gao, Liu, Zhuang, and Weinberger, Kilian Q · 2016
Closest in time.
Deep learning without poor local minima
Kawaguchi, Kenji · 2016
Closest in time.
On the depth of deep neural networks: A theoretical view
Sun, Shizhao, Chen, Wei, Wang, Liwei, Liu, Xiaoguang, and Liu, Tie-Yan · 2016
Closest in time.
Benefits of depth in neural networks
Telgarsky, Matus · 2016
Closest in time.
Architectural complexity measures of recurrent neural networks
Zhang, Saizheng, Wu, Yuhuai, Che, Tong, Lin, Zhouhan, Memisevic, Roland, Salakhutdinov, Ruslan, and Bengio, Yoshua · 2016
Closest in time.