Fetching the paper…
Reading the bibliography…
Successive linear transforms followed by nonlinear "activation" functions can approximate nonlinear functions to arbitrary precision given sufficient layers.
M. Leshno, V. Y. Lin, A. Pinkus, and S. Schocken, “Multilayer feedforward networks with a nonpolynomial activation function can approximate any function,” Neural Networks
1993
Earlier work this paper cites.
A. Krizhevsky, “Learning multiple layers of features from tiny images,” tech. rep., 2009
2009
Earlier work this paper cites.
R. Arora, A. Basu, P. Mianjy, and A. Mukherjee, “Understanding deep neural networks with rectified linear units,” 2016
2016
Cited alongside, same era.
B. Xu, R. Huang, and M. Li, “Revise saturated activation functions,” 2016
2016
Cited alongside, same era.
P. Ramachandran, B. Zoph, and Q. V. Le, “Searching for activation functions,” 2017
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…