Fetching the paper…
Reading the bibliography…
We provide novel guaranteed approaches for training feedforward neural networks with sparse connectivity.
A method of solving a convex programming problem with convergence rate o (1/k2)
Yurii Nesterov · 1983
Earlier work this paper cites.
Approximate computation of expectations
Charles Stein · 1986
Earlier work this paper cites.
On principal hessian directions for data visualization and dimension reduction: another application of stein’s lemma
Ker-Chau Li · 1992
Earlier work this paper cites.
Principal hessian directions revisited
R Dennis Cook · 1998
Earlier work this paper cites.
Exploiting generative models in discriminative classifiers
Tommi Jaakkola, David Haussler, et al · 1999
Earlier work this paper cites.
An iterative thresholding algorithm for linear inverse problems with a sparsity constraint
Ingrid Daubechies, Michel Defrise, and Christine De Mol · 2004
Earlier work this paper cites.
Use of exchangeable pairs in the analysis of simulations
Charles Stein, Persi Diaconis, Susan Holmes, Gesine Reinert, et al · 2004
Earlier work this paper cites.
Convex neural networks
Yoshua Bengio, Nicolas L Roux, Pascal Vincent, Olivier Delalleau, and Patrice Marcotte · 2005
Earlier work this paper cites.
Estimation of non-normalized statistical models by score matching
Aapo Hyvärinen · 2005
Earlier work this paper cites.
Pattern recognition and machine learning , volume 1
Christopher M Bishop et al · 2006
Earlier work this paper cites.
Gradient projection for sparse reconstruction: Application to compressed sensing and other inverse problems
Mário AT Figueiredo, Robert D Nowak, and Stephen J Wright · 2007
Cited alongside, same era.
An interior-point method for large-scale l 1-regularized least squares
Seung-Jean Kim, Kwangmoo Koh, Michael Lustig, Stephen Boyd, and Dimitry Gorinevsky · 2007
Cited alongside, same era.
Gradient methods for minimizing composite objective function, 2007
Yurii Nesterov et al · 2007
Cited alongside, same era.
Multi-label classification: An overview
Grigorios Tsoumakas and Ioannis Katakis · 2007
Cited alongside, same era.
On stein identity, chernoff inequality, and orthogonal polynomials
Zhengyuan Wei, Xinsheng Zhang, and Taifu Li · 2010
Cited alongside, same era.
On autoencoders and score matching for energy based models
Kevin Swersky, David Buchman, Nando D Freitas, Benjamin M Marlin, et al · 2011
Exact recovery of sparsely-used dictionaries
Daniel A Spielman, Huan Wang, and John Wright · 2012
Later among the works it cites.
Provable bounds for learning some deep representations
Sanjeev Arora, Aditya Bhaskara, Rong Ge, and Tengyu Ma · 2013
Later among the works it cites.
Low-rank approximations for conditional feedforward computation
Andrew Davis and Itamar Arel · 2013
Later among the works it cites.
Integration by parts and representation of information functionals
Ivan Nourdin, Giovanni Peccati, and Yvik Swan · 2013
Later among the works it cites.
Learning mixtures of linear classifiers
Yuekai Sun, Stratis Ioannidis, and Andrea Montanari · 2013
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
What regularized auto-encoders learn from the data generating distribution
Guillaume Alain and Yoshua Bengio · 2012
Cited alongside, same era.
Learning Topic Models and Latent Bayesian Networks Under Expansion Constraints
A. Anandkumar, D. Hsu, and A. Javanmard S. M. Kakade · 2012
Cited alongside, same era.
Hypercontractivity, sum-of-squares proofs, and their applications
Boaz Barak, Fernando GSL Brandao, Aram W Harrow, Jonathan Kelner, David Steurer, and Yuan Zhou · 2012
Cited alongside, same era.
Contributions to the mathematical theory of evolution
K. Pearson
Cited in the paper.
Later among the works it cites.
Sparse activity and sparse connectivity in supervised learning
Markus Thom and Günther Palm · 2013
Later among the works it cites.
Tensor decompositions for learning latent variable models
A. Anandkumar, R. Ge, D. Hsu, S. M. Kakade, and M. Telgarsky · 2014
Closest in time.
Clustering via mode seeking by direct estimation of the gradient of a log-density
Hiroaki Sasaki, Aapo Hyvärinen, and Masashi Sugiyama · 2014
Closest in time.
Dictionary learning with few samples and matrix concentration
Kyle Luh and Van Vu · 2015
Closest in time.