Fetching the paper…
Reading the bibliography…
We give the first dimension-efficient algorithms for learning Rectified Linear Units (ReLUs), which are functions of the form $\mathbf{x} \mapsto \max(0, \mathbf{w} \cdot \mathbf{x})$ with $\mathbf{w} \in \mathbb{S}^{n-1}$.
Functions of positive and negative type, and their connection with the theory of integral equations
James Mercer · 1909
Earlier work this paper cites.
Rational approximation to | x | |x|
D. J. Newman · 1964
Earlier work this paper cites.
Multivariate adaptive regression splines
Jerome H. Friedman · 1991
Earlier work this paper cites.
Probability in Banach Spaces: Isoperimetry and Processes
Michel Ledoux and Michel Talagrand · 1991
Earlier work this paper cites.
Decision theoretic generalizations of the pac model for neural net and other learning applications
David Haussler · 1992
Earlier work this paper cites.
Toward efficient agnostic learning
Michael J. Kearns, Robert E. Schapire, and Linda M. Sellie · 1994
Earlier work this paper cites.
An introduction to support vector machines and other kernel-based learning methods
Nello Cristianini and John Shawe-Taylor · 2000
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L. Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Noise-tolerant learning, the parity problem, and the statistical query model
Blum, Kalai, and Wasserman · 2003
Earlier work this paper cites.
Kernel methods in machine learning
Thomas Hofmann, Bernhard Schölkopf, and Alexander J Smola · 2008
Earlier work this paper cites.
On the complexity of linear prediction: Risk bounds, margin bounds, and regularization
Sham M. Kakade, Karthik Sridharan, and Ambuj Tewari · 2008
Earlier work this paper cites.
Agnostically learning halfspaces
Adam Tauman Kalai, Adam R. Klivans, Yishay Mansour, and Rocco A. Servedio · 2008
Earlier work this paper cites.
On agnostic learning of parities, monomials, and halfspaces
V. Feldman, P. Gopalan, S. Khot, and A. K. Ponnuswami · 2009
Cited alongside, same era.
Cryptographic hardness for learning intersections of halfspaces
A. R. Klivans and A. A. Sherstov · 2009
Cited alongside, same era.
Convex piecewise-linear fitting
Alessandro Magnani and Stephen P. Boyd · 2009
Cited alongside, same era.
Bounded independence fools degree-2 threshold functions
Ilias Diakonikolas, Daniel M. Kane, and Jelani Nelson · 2010
Cited alongside, same era.
Learning kernel-based halfspaces with the 0-1 loss
Shai Shalev-Shwartz, Ohad Shamir, and Karthik Sridharan · 2011
Cited alongside, same era.
Reliable agnostic learning
Adam Tauman Kalai, Varun Kanade, and Yishay Mansour · 2012
Cited alongside, same era.
Agnostic learning of disjunctions on symmetric distributions
Vitaly Feldman and Pravesh Kothari · 2015
Later among the works it cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Later among the works it cites.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Later among the works it cites.
High-Dimensional Statistics
Phillippe Rigollet · 2015
Later among the works it cites.
Finding correlations in subquadratic time, with applications to learning parities and the closest pair problem
Gregory Valiant · 2015
Later among the works it cites.
Understanding deep neural networks with rectified linear units, 2016
Raman Arora, Amitabh Basu, Poorya Mianjy, and Anribit Mukherjee · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Making polynomials robust to noise
Alexander A. Sherstov · 2012
Cited alongside, same era.
Rectifier nonlinearities improve neural network acoustic models
Andrew L Maas, Awni Y Hannun, and Andrew Y Ng · 2013
Cited alongside, same era.
Learning sparse polynomial functions
Alexandr Andoni, Rina Panigrahy, Gregory Valiant, and Li Zhang · 2014
Cited alongside, same era.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2014
Cited alongside, same era.
Embedding hard learning problems into gaussian space
Adam Klivans and Pravesh Kothari · 2014
Cited alongside, same era.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Cited alongside, same era.
Closest in time.
Complexity theoretic limitations on learning halfspaces
Amit Daniely · 2016
Closest in time.
The power of depth for feedforward neural networks
Ronen Eldan and Ohad Shamir · 2016
Closest in time.
Approve or Reject: Can You Moderate Five New York Times Comments?
Bassey Etim · 2016
Closest in time.
Multinomial theorem — Wikipedia, the free encyclopedia, 2016
Wikipedia · 2016
Closest in time.
Polynomial kernel — Wikipedia, the free encyclopedia, 2016
Wikipedia · 2016
Closest in time.
ℓ 1 \ell_{1} networks are improperly learnable in polynomial-time
Yuchen Zhang, Jason Lee, and Michael Jordan · 2016
Closest in time.