Fetching the paper…
Reading the bibliography…
The step function is one of the simplest and most natural activation functions for deep neural networks (DNNs).
“A logical calculus of the ideas immanent in nervous activity,”
W. S. Mcculloch and W. H. Pitts, · 1943
Earlier work this paper cites.
“Computing a trust region step,”
J. J. More and D. C. Sorensen, · 1983
Earlier work this paper cites.
“Learning representations by back-propagating errors,”
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, · 1986
Earlier work this paper cites.
“The influence of the sigmoid function parameters on the speed of backpropagation learning,”
J. Han and C. Moraga, · 1995
Earlier work this paper cites.
“Variational analysis,”
R. T. Rockafellar and R. J. Wets, · 2009
Earlier work this paper cites.
“Rectified linear units improve restricted boltzmann machines,”
V. Nair and G. E. Hinton, · 2010
Earlier work this paper cites.
“MNIST handwritten digit database,” 2010
Y. LeCun, and C. Cortes · 2010
Earlier work this paper cites.
“Adaptive subgradient methods for online learning and stochastic optimization,”
J. Duchi, E. Hazan, and Y. Singer, · 2011
Earlier work this paper cites.
“Neural networks for machine learning lecture 6a overview of mini-batch gradient descent,”
G. Hinton, N. Srivastava and K. Swersky · 2012
Earlier work this paper cites.
“Sensitivity-based adaptive learning rules for binary feedforward neural networks,”
S. M. Zhong, X. Q. Zeng, S. L. Wu and L. X. Han · 2012
Earlier work this paper cites.
“Lecture 6.5 - RMSProp, COURSERA: Neural networks for machine learning,”
T. Tieleman and G. Hinton, · 2012
Earlier work this paper cites.
“Adadelta: an adaptive learning rate method,”
M. D. Zeiler, · 2012
Earlier work this paper cites.
“Rectifier nonlinearities improve neural network acoustic models,”
A. L. Maas, A. Y. Hannun, and A. Y. Ng, · 2013
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
D. P. Kingma and J. Ba, · 2014
Earlier work this paper cites.
“The CIFAR-10 dataset,”
A. Krizhevsky, V. Nair, and G. Hinton · 2014
Earlier work this paper cites.
“Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,”
K. He, X. Zhang, S. Ren, and J. Sun, · 2015
Earlier work this paper cites.
“Fast and accurate deep network learning by exponential linear units (ELUs),”
D. A. Clevert, T. Unterthiner, and S. Hochreiter, · 2015
Earlier work this paper cites.
“Binaryconnect: Training deep neural networks with binary weights during propagations”,
M. Courbariaux, Y. Bengio and J-P. David, · 2015
Earlier work this paper cites.
“BinaryConnect: Training deep neural networks with binary weights during propagations,”
M. Courbariaux and Y. S. Bengio, · 2015
Earlier work this paper cites.
“Deep learning,”
I. Goodfellow, Y. Bengio, and A. Courville, · 2016
Earlier work this paper cites.
“Incorporating Nesterov momentum into Adam,”
T. Dozat, · 2016
Earlier work this paper cites.
“An overview of gradient descent optimization algorithms,”
S. Ruder, · 2016
Cited alongside, same era.
“Training neural networks without gradients: A scalable ADMM approach,”
G. Taylor, R. Burmeister, Z. Xu, B. Singh, A. Patel, and T. Goldstein, · 2016
Cited alongside, same era.
“Binarised neural networks: Training deep neural networks with weights and activations constrained to +1 or -1,”
M. Courbariaux, I. Hubara, D. Soudry and Y. S. Bengio, · 2016
Cited alongside, same era.
“Self-normalizing neural networks,”
G. Klambauer, T. Unterthiner, A. Mayr, and S. Hochreiter, · 2017
Cited alongside, same era.
“Swish: a self-gated activation function,”
P. Ramachandran, B. Zoph, and Q. V. Le, · 2017
Cited alongside, same era.
“Channel pruning for deep neural networks via a relaxed group-wise splitting method,”
B. Yang, J. Lyu, S. Zhang, Y. Qi, and J. Xin, · 2019
Later among the works it cites.
“Global convergence of block coordinate descent in deep learning,”
J. S. Zeng, T. T. Lau, S. B. Lin, and Y. Yao, · 2019
Later among the works it cites.
“Optimization problems involving group sparsity terms,”
A. Beck and N. Hallak, · 2019
Later among the works it cites.
“Training binary neural networks with real-to-binary convolutions”,
B. Martinez, J. Yang, A. Bulat, G. Tzimiropoulos, · 2020
Later among the works it cites.
“Training binary neural networks with real-to-binary convolutions. ”
B. Martinez, J. Yang, A. Bulat, and G. Tzimiropoulos · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“How to train a compact binary neural network with high accuracy”,
W. Tang, G. Hua, L. Wang, · 2017
Cited alongside, same era.
“Convergent block coordinate descent for training Tikhonov regularized deep neural networks,”
Z. Zhang and M. Brand, · 2017
Cited alongside, same era.
“Group sparse optimization via l p , q l_{p,q} regularization,”
Y. Hu, C. Li, K. Meng, J. Qin, and X. Yang, · 2017
Cited alongside, same era.
“Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms,”
H. Xiao, K. Rasul, and R. Vollgraf, · 2017
Cited alongside, same era.
“An empirical study of binary neural networks’ optimisation”,
M. Alizadeh, J. Fern andez-Marques, N. Lane and Y. Gal, · 2018
Cited alongside, same era.
“On the convergence of Adam and beyond,”
S. J. Reddi, S. Kale, and S. Kumar, · 2018
Cited alongside, same era.
“A proximal block coordinate descent algorithm for deep neural network training,”
T. T. Lau, J. Zeng, B. Wu, and Y. Yao, · 2018
Cited alongside, same era.
T. Dinh, B. Wang, A. L. Bertozzi, and S. J. Osher, · 2020
Later among the works it cites.
“Machine learning”
Z. H. Zhou, · 2021
Later among the works it cites.
“High-capacity expert binary networks”,
A. Bulat, B. Martinez and G. Tzimiropoulos, · 2021
Later among the works it cites.
“Distillation-guided residual learning for binary convolutional neural networks.”
J. M. Ye, J. D. Wang, and S. L. Zhang · 2021
Later among the works it cites.
“Modulated convolutional networks.”
B. C. Zhang, R. Q. Wang, X.D. Wang, J.G. Han, and R. R. Ji · 2021
Later among the works it cites.
“Support vector machine classifier via L 0 / 1 L_{0/1} soft-margin loss,”
H. J. Wang, Y. H. Shao, S. L. Zhou, C. Zhang, and N. H. Xiu, · 2021
Later among the works it cites.
“Quadratic convergence of smoothing Newton’s method for 0/1 loss optimization,”
S. L. Zhou, L. L. Pan, N. H. Xiu, and H. D. Qi, · 2021
Later among the works it cites.
“Recursion Newton-like algorithm for ℓ 2 , 0 \ell_{2,0} -ReLU deep neural networks,”
H. Zhang, Z. P. Yuan, and N. H. Xiu, · 2021
Later among the works it cites.
T. Hoefler, D. Alistarh, T. Ben-Nun, N. Dryden, and A. Peste, · 2021
Later among the works it cites.
“Newton method for ℓ 0 \ell_{0} -regularized optimization,”
S. L. Zhou, L. L. Pan, and N. H. Xiu, · 2021
Later among the works it cites.
“Toward pixel-level precision for binary super-resolution with mixed binary representation.”
X. R. Jiang, N. Wang, J. Xin, K. Li, X. Yang, J. Li, X. Gao · 2022
Closest in time.
“Local means binary networks for image super-resolution.”
K. Li, N. N. Wang, J. W. Xin, X. R. Jiang, J. Li, X. B. Gao, K. Han, and Y. H. Wang · 2022
Closest in time.
“Toward accurate binarized neural networks with sparsity for mobile application.”
P. S. Wang, X. Y. He, and J. Cheng · 2022
Closest in time.
“Computing one-bit compressive sensing via double-sparsity constrained optimization,”
S. L. Zhou, Z. Y. Luo, N H. Xiu, and G. Y. Li, · 2022
Closest in time.
“A comprehensive review of binary neural network,”
C Yuan and S. Agaian, · 2023
Closest in time.