Fetching the paper…
Reading the bibliography…
We present experiments demonstrating that some other form of capacity control, different from network size, plays a central role in learning multilayer feed-forward networks.
Cryptographic limitations on learning boolean formulae and finite automata
Kearns, Michael and Valiant, Leslie · 1994
Earlier work this paper cites.
Maximum-margin matrix factorization
Srebro, Nathan, Rennie, Jason, and Jaakkola, Tommi S · 1994
Earlier work this paper cites.
Neural network learning: Theoretical foundations
Anthony, Martin and Bartlett, Peter L · 1999
Earlier work this paper cites.
A rank minimization heuristic with application to minimum order system approximation
Fazel, Maryam, Hindi, Haitham, and Boyd, Stephen P · 2001
Earlier work this paper cites.
Weighted low-rank approximations
Srebro, Nathan and Jaakkola, Tommi S · 2003
Earlier work this paper cites.
Convex neural networks
Bengio, Yoshua, Roux, Nicolas L., Vincent, Pascal, Delalleau, Olivier, and Marcotte, Patrice · 2005
Earlier work this paper cites.
Fast maximum margin matrix factorization for collaborative prediction
Rennie, Jasson DM and Srebro, Nathan · 2005
Cited alongside, same era.
Computational enhancements in low-rank semidefinite programming
Burer, Samuel and Choi, Changhui · 2006
Cited alongside, same era.
Cryptographic hardness for learning intersections of halfspaces
Sherstov, Adam R Klivansand Alexander A · 2006
Cited alongside, same era.
Introduction to the Theory of Computation
Sipser, Michael · 2006
Cited alongside, same era.
Multi-task feature learning
Argyriou, Andreas, Evgeniou, Theodoros, and Pontil, Massimiliano · 2007
Cited alongside, same era.
ℓ 1 \ell_{1} regularization in infinite dimensional feature spaces
Rosset, Saharon, Swirszcz, Grzegorz, and Srebro, Nathan · 2007
Cited alongside, same era.
Kernel methods for deep learning
Cho, Youngmin and Saul, Lawrence K · 2009
Later among the works it cites.
Collaborative filtering in a non-uniform world: Learning with the weighted trace norm
Srebro, Nathan and Salakhutdinov, Ruslan · 2010
Later among the works it cites.
Breaking the curse of dimensionality with convex neural networks
Bach, Francis · 2014
Closest in time.
From average case complexity to improper learning complexity
Daniely, Amit, Linial, Nati, and Shalev-Shwartz, Shai · 2014
Closest in time.
On the computational efficiency of training neural networks
Livni, Roi, Shalev-Shwartz, Shai, and Shamir, Ohad · 2014
Closest in time.
How to scale up kernel methods to be as good as deep neural nets
Lu, Zhiyun, May, Avner, Liu, Kuan, Garakani, Alireza Bagheri, Guo, Dong, Bellet, Aurlien, Fan, Linxi, Collins, Michael, Kingsbury, Brian, Picheny, Michael, and Sha, Fei · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Closest in time.