Fetching the paper…
Reading the bibliography…
The goal of this paper is to characterize function distributions that deep learning can or cannot learn in poly-time.
Johan Håstad, Computational limitations of small-depth circuits , MIT Press, Cambridge, MA, USA, 1987
1987
Earlier work this paper cites.
Marvin Minsky and Seymour Papert, Perceptrons - an introduction to computational geometry , MIT Press, 1987
1987
Earlier work this paper cites.
B. Chor and O. Goldreich, Unbiased bits from sources of weak randomness and probabilistic communication complexity , SIAM Journal on Computing 17
1988
Earlier work this paper cites.
Avrim Blum, Merrick Furst, Jeffrey Jackson, Michael Kearns, Yishay Mansour, and Steven Rudich, Weakly learning dnf and characterizing statistical query learning using fourier analysis , Proceedings of the Twenty-sixth Annual ACM Symposium on Theory of Computing (New York, NY, USA), STOC ’94, ACM, 1994, pp. 253–262
1994
Earlier work this paper cites.
Ian Parberry, Circuit complexity and neural networks , MIT Press, Cambridge, MA, USA, 1994
1994
Earlier work this paper cites.
Eric Allender, Circuit complexity before the dawn of the new millennium , Foundations of Software Technology and Theoretical Computer Science (Berlin, Heidelberg) (V. Chandru and V. Vinay, eds.), Springer Berlin Heidelberg, 1996, pp. 1–18
1996
Earlier work this paper cites.
Michael Kearns, Efficient noise-tolerant learning from statistical queries , J. ACM 45
1998
Earlier work this paper cites.
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner, Gradient-based learning applied to document recognition , Proceedings of the IEEE 86
1998
Earlier work this paper cites.
Avrim Blum, Adam Kalai, and Hal Wasserman, Noise-tolerant learning, the parity problem, and the statistical query model , J. ACM 50
2003
Earlier work this paper cites.
Oded Regev, On lattices, learning with errors, random linear codes, and cryptography , Proceedings of the Thirty-seventh Annual ACM Symposium on Theory of Computing (New York, NY, USA), STOC ’05, ACM, 2005, pp. 84–93
2005
Earlier work this paper cites.
Michael Sipser, Introduction to the theory of computation , second ed., Course Technology, 2006
2006
Earlier work this paper cites.
Martin Anthony and Peter L. Bartlett, Neural network learning: Theoretical foundations , 1st ed., Cambridge University Press, New York, NY, USA, 2009
2009
Earlier work this paper cites.
Adam R. Klivans and Alexander A. Sherstov, Cryptographic hardness for learning intersections of halfspaces , J. Comput. Syst. Sci. 75
2009
Earlier work this paper cites.
Max Welling and Yee Whye Teh, Bayesian learning via stochastic gradient langevin dynamics , Proceedings of the 28th International Conference on International Conference on Machine Learning (USA), ICML’11, Omnipress, 2011, pp. 681–688
2011
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury, Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups , IEEE Signal Processing Magazine 29
2012
Earlier work this paper cites.
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton, Imagenet classification with deep convolutional neural networks , Advances in Neural Information Processing Systems 25 (F. Pereira, C. J. C. Burges, L. Bottou, and K. Q. Weinberger, eds.), Curran Associates, Inc., 2012, pp. 1097–1105
2012
Earlier work this paper cites.
Shai Shalev-Shwartz and Shai Ben-David, Understanding machine learning: From theory to algorithms , Cambridge University Press, New York, NY, USA, 2014
2014
Cited alongside, same era.
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan, Escaping from saddle points — online stochastic gradient for tensor decomposition , Proceedings of The 28th Conference on Learning Theory (Paris, France) (Peter Grünwald, Elad Hazan, and Satyen Kale, eds.), Proceedings of Machine Learning Research, vol. 40, PMLR, 03–06 Jul 2015, pp. 797–842
2015
Cited alongside, same era.
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, Delving deep into rectifiers: Surpassing human-level performance on imagenet classification , 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile, December 7-13, 2015, 2015, pp. 1026–1034
2015
Cited alongside, same era.
Yann Lecun, Yoshua Bengio, and Geoffrey Hinton, Deep learning , Nature 521
2015
Cited alongside, same era.
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer, Automatic differentiation in pytorch , NIPS-W, 2017
2017
Later among the works it cites.
Ioannis Panageas and Georgios Piliouras, Gradient descent only converges to minimizers: Non-isolated critical points and invariant regions , ITCS, 2017
2017
Later among the works it cites.
Ran Raz, A time-space lower bound for a large class of learning problems , 2017 IEEE 58th Annual Symposium on Foundations of Computer Science (FOCS) (2017), 732–742
2017
Later among the works it cites.
M. Raginsky, A. Rakhlin, and M. Telgarsky, Non-convex learning via Stochastic Gradient Langevin Dynamics: a nonasymptotic analysis , ArXiv e-prints (2017)
2017
Later among the works it cites.
S. Shalev-Shwartz, O. Shamir, and S. Shammah, Failures of Gradient-Based Deep Learning , ArXiv e-prints (2017)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jacob Steinhardt, Gregory Valiant, and Stefan Wager, Memory, communication, and statistical queries , Electronic Colloquium on Computational Complexity, 2015
2015
Cited alongside, same era.
Amit Daniely and Shai Shalev-Shwartz, Complexity theoretic limitations on learning dnf’s , 29th Annual Conference on Learning Theory (Columbia University, New York, New York, USA) (Vitaly Feldman, Alexander Rakhlin, and Ohad Shamir, eds.), Proceedings of Machine Learning Research, vol. 49, PMLR, 23–26 Jun 2016, pp. 815–830
2016
Cited alongside, same era.
Ian Goodfellow, Yoshua Bengio, and Aaron Courville, Deep learning , MIT Press, 2016, http://www.deeplearningbook.org
2016
Cited alongside, same era.
Moritz Hardt, Benjamin Recht, and Yoram Singer, Train faster, generalize better: Stability of stochastic gradient descent , Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48, ICML’16, JMLR.org, 2016, pp. 1225–1234
2016
Cited alongside, same era.
R. Raz, Fast Learning Requires Good Memory: A Time-Space Lower Bound for Parity Learning , ArXiv e-prints (2016)
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Later among the works it cites.
2017
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
Sumegha Garg, Ran Raz, and Avishay Tal, Extractor-based time-space lower bounds for learning , Proceedings of the 50th Annual ACM SIGACT Symposium on Theory of Computing (New York, NY, USA), STOC 2018, ACM, 2018, pp. 990–1002
2018
Later among the works it cites.
R. Kleinberg, Y. Li, and Y. Yuan, An Alternative View: When Does SGD Escape Local Minima? , ArXiv e-prints (2018)
2018
Later among the works it cites.
Ohad Shamir, Distribution-specific hardness of learning neural networks , Journal of Machine Learning Research 19
2018
Later among the works it cites.
2018
Later among the works it cites.
E. Bamas, Semester Project Report,
2019
Later among the works it cites.
E. Boix, MDS Internal Report,
2019
Later among the works it cites.
L. Bottou, Personal communication, 2019
2019
Later among the works it cites.