Fetching the paper…
Reading the bibliography…
In this paper, we consider the joint task of simultaneously optimizing (i) the weights of a deep neural network, (ii) the number of neurons for each hidden layer, and (iii) the subset of active input features (i.e., feature selection).
Y. LeCun, J. S. Denker, S. A. Solla, R. E. Howard, and L. D. Jackel, “Optimal brain damage.” in Advances in Neural Information Processing Systems , vol. 2, 1989, pp. 598–605
1989
Earlier work this paper cites.
J. Moody, S. Hanson, A. Krogh, and J. A. Hertz, “A simple weight decay can improve generalization,” Advances in Neural Information Processing Systems , vol. 4, pp. 950–957, 1995
1995
Earlier work this paper cites.
R. Tibshirani, “Regression shrinkage and selection via the lasso,” Journal of the Royal Statistical Society. Series B (Methodological) , pp. 267–288, 1996
1996
Earlier work this paper cites.
F. Alimoglu and E. Alpaydin, “Methods of combining multiple classifiers based on different representations for pen-based handwritten digit recognition,” in Proceedings of the Fifth Turkish Artificial Intelligence and Artificial Neural Networks Symposium (TAINN . Citeseer, 1996
1996
Earlier work this paper cites.
J. A. Blackard and D. J. Dean, “Comparative accuracies of artificial neural networks and discriminant analysis in predicting forest cover types from cartographic variables,” Computers and electronics in agriculture , vol. 24, no. 3, pp. 131–151, 1999
1999
Earlier work this paper cites.
A. Ben-Tal and A. Nemirovski, Lectures on modern convex optimization: analysis, algorithms, and engineering applications . SIAM, 2001, vol. 2
2001
Earlier work this paper cites.
K. Suzuki, I. Horiba, and N. Sugie, “A simple neural network pruning algorithm with application to filter synthesis,” Neural Processing Letters , vol. 13, no. 1, pp. 43–53, 2001
2001
Earlier work this paper cites.
N. Kwak and C.-H. Choi, “Input feature selection for classification problems,” IEEE Transactions on Neural Networks , vol. 13, no. 1, pp. 143–159, 2002
2002
Earlier work this paper cites.
I. Guyon and A. Elisseeff, “An introduction to variable and feature selection,” Journal of machine learning research , vol. 3, no. Mar, pp. 1157–1182, 2003
2003
Earlier work this paper cites.
K. Pelckmans, J. A. Suykens, and B. De Moor, “Morozov, ivanov and tikhonov regularization based ls-svms,” in International Conference on Neural Information Processing . Springer, 2004, pp. 1216–1222
2004
Earlier work this paper cites.
H. Zou and T. Hastie, “Regularization and variable selection via the elastic net,” Journal of the Royal Statistical Society. Series B (Statistical Methodology) , vol. 67, no. 2, pp. 301–320, 2005
2005
Earlier work this paper cites.
M. Yuan and Y. Lin, “Model selection and estimation in regression with grouped variables,” Journal of the Royal Statistical Society. Series B (Statistical Methodology) , vol. 68, no. 1, pp. 49–67, 2006
2006
Earlier work this paper cites.
E. J. Candès and M. B. Wakin, “An introduction to compressive sampling,” IEEE Signal Processing Magazine , vol. 25, no. 2, pp. 21–30, 2008
2008
Earlier work this paper cites.
F. R. Bach, “Consistency of the group lasso and multiple kernel learning,” Journal of Machine Learning Research , vol. 9, no. Jun, pp. 1179–1225, 2008
2008
Earlier work this paper cites.
J. Liu, S. Ji, and J. Ye, “Multi-task feature learning via efficient l 2, 1-norm minimization,” in Proceedings of the Twenty-Fifth Conference on Uncertainty in Artificial Intelligence . AUAI Press, 2009, pp. 339–348
2009
Earlier work this paper cites.
S. S. Haykin, Neural networks and learning machines , 3rd ed. Pearson, 2009
2009
Earlier work this paper cites.
M. Schmidt, “Graphical model structure learning with l1-regularization,” Ph.D. dissertation, University of British Columbia (Vancouver), 2010
2010
Earlier work this paper cites.
2010
Cited alongside, same era.
J. Bergstra, O. Breuleux, F. Bastien, P. Lamblin, R. Pascanu, G. Desjardins, J. Turian, D. Warde-Farley, and Y. Bengio, “Theano: A cpu and gpu math compiler in python,” in Proc. 9th Python in Science Conf , 2010, pp. 1–7
2010
Cited alongside, same era.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in International Conference on Artificial Intelligence and Statistics , 2010, pp. 249–256
2010
Cited alongside, same era.
B. Recht, C. Re, S. Wright, and F. Niu, “Hogwild: A lock-free approach to parallelizing stochastic gradient descent,” in Advances in Neural Information Processing Systems , 2011, pp. 693–701
2011
Cited alongside, same era.
2014
Later among the works it cites.
2014
Later among the works it cites.
G. Scutari, F. Facchinei, P. Song, D. P. Palomar, and J.-S. Pang, “Decomposition by partial linearization: Parallel optimization of multi-agent systems,” IEEE Transactions on Signal Processing , vol. 62, no. 3, pp. 641–656, 2014
2014
Later among the works it cites.
J. Schmidhuber, “Deep learning in neural networks: An overview,” Neural Networks , vol. 61, pp. 85–117, 2015
2015
Later among the works it cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature , vol. 521, no. 7553, pp. 436–444, 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Jenatton, J.-Y. Audibert, and F. Bach, “Structured variable selection with sparsity-inducing norms,” Journal of Machine Learning Research , vol. 12, no. Oct, pp. 2777–2824, 2011
2011
Cited alongside, same era.
X. Glorot, A. Bordes, and Y. Bengio, “Deep sparse rectifier neural networks,” in International Conference on Artificial Intelligence and Statistics , 2011, pp. 315–323
2011
Cited alongside, same era.
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay, “Scikit-learn: Machine learning in Python,” Journal of Machine Learning Research , vol. 12, pp. 2825–2830, 2011
2011
Cited alongside, same era.
P. Domingos, “A few useful things to know about machine learning,” Communications of the ACM , vol. 55, no. 10, pp. 78–87, 2012
2012
Cited alongside, same era.
F. Bach, R. Jenatton, J. Mairal, and G. Obozinski, “Optimization with sparsity-inducing penalties,” Foundations and Trends® in Machine Learning , vol. 4, no. 1, pp. 1–106, 2012
2012
Cited alongside, same era.
Y. Bengio, “Practical recommendations for gradient-based training of deep architectures,” in Neural Networks: Tricks of the Trade . Springer, 2012, pp. 437–478
2012
Cited alongside, same era.
L. Deng, “The mnist database of handwritten digit images for machine learning research,” IEEE Signal Processing Magazine , vol. 29, no. 6, pp. 141–142, 2012
2012
Cited alongside, same era.
M. Denil, B. Shakibi, L. Dinh, and N. de Freitas, “Predicting parameters in deep learning,” in Advances in Neural Information Processing Systems , 2013, pp. 2148–2156
2013
Cited alongside, same era.
2015
Later among the works it cites.
2015
Later among the works it cites.
M. Courbariaux, Y. Bengio, and J.-P. David, “Binaryconnect: Training deep neural networks with binary weights during propagations,” in Advances in Neural Information Processing Systems , 2015, pp. 3123–3131
2015
Later among the works it cites.
S. Han, J. Pool, J. Tran, and W. Dally, “Learning both weights and connections for efficient neural network,” in Advances in Neural Information Processing Systems , 2015, pp. 1135–1143
2015
Later among the works it cites.
2015
Later among the works it cites.
S. Gupta, A. Agrawal, K. Gopalakrishnan, and P. Narayanan, “Deep learning with limited numerical precision,” in Proceedings of The 32nd International Conference on Machine Learning , 2015, pp. 1737–1746
2015
Later among the works it cites.
W. Chen, J. Wilson, S. Tyree, K. Weinberger, and Y. Chen, “Compressing neural networks with the hashing trick,” in Proceedings of The 32nd International Conference on Machine Learning , 2015, pp. 2285–2294
2015
Later among the works it cites.
L. Zhao, Q. Hu, and W. Wang, “Heterogeneous feature selection with multi-modal deep neural networks and sparse group lasso,” IEEE Transactions on Multimedia , vol. 17, no. 11, pp. 1936–1948, 2015
2015
Later among the works it cites.
P. Ochs, A. Dosovitskiy, T. Brox, and T. Pock, “On iteratively reweighted algorithms for nonsmooth nonconvex optimization in computer vision,” SIAM Journal on Imaging Sciences , vol. 8, no. 1, pp. 331–372, 2015
2015
Later among the works it cites.
F. M. Bianchi, S. Scardapane, A. Uncini, A. Rizzi, and A. Sadeghian, “Prediction of telephone calls load using echo state network with exogenous variables,” Neural Networks , vol. 71, pp. 204–213, 2015
2015
Later among the works it cites.
2016
Closest in time.
2016
Closest in time.
S. Scardapane, R. Fierimonte, P. Di Lorenzo, M. Panella, and A. Uncini, “Distributed semi-supervised support vector machines,” Neural Networks , vol. 80, pp. 43–52, 2016
2016
Closest in time.