Fetching the paper…
Reading the bibliography…
As neural networks grow deeper and wider, learning networks with hard-threshold activations is becoming increasingly important, both for network quantization, which can drastically reduce time and energy requirements, and for creating large integrated systems of deep networks, which may have non-differentiable components and must avoid vanishing and exploding gradients for effective learning.
The perceptron: A probabilistic model for information storage and organization in the brain
Frank Rosenblatt · 1958
Earlier work this paper cites.
On convergence proofs on perceptrons
A. B. J. Novikoff · 1962
Earlier work this paper cites.
Perceptrons: an introduction to computational geometry
Marvin L. Minsky and Seymour Papert · 1969
Earlier work this paper cites.
Learning Process in an Asymmetric Threshold Network
Yann LeCun · 1986
Earlier work this paper cites.
Learining Internal Representations by Error Propagation
David E. Rumelhart, Geoffrey E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Modeles connexionnistes de l’apprentissage (connectionist learning models)
Yann LeCun · 1987
Earlier work this paper cites.
MADALINE RULE II: A training algorithm for neural networks
Rodney Winter and Bernard Widrow · 1988
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Robust Truncated Hinge Loss Support Vector Machines
Yichao Wu and Yufeng Liu · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Coursera Lectures: Neural networks for machine learning, 2012
Geoffrey E. Hinton · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Earlier work this paper cites.
Efficient BackProp
Yann LeCun, Léon Bottou, Genevieve B. Orr, and Klaus-Robert Müller · 2012
Cited alongside, same era.
Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville · 2013
Cited alongside, same era.
How Auto-Encoders Could Provide Credit Assignment in Deep Networks via Target Propagation
Yoshua Bengio · 2014
Cited alongside, same era.
Distributed optimization of deeply nested systems
Miguel Á. Carreira-Perpiñán and Weiran Wang · 2014
Cited alongside, same era.
Expectation Backpropagation: parameter-free training of multilayer neural networks with real and discrete weights
Daniel Soudry, Itay Hubara, and Ron Meir · 2014
Cited alongside, same era.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Later among the works it cites.
Fixed Point Quantization of Deep Convolutional Networks
Darryl D. Lin and Sachin S. Talathi · 2016
Later among the works it cites.
Overcoming Challenges in Fixed Point Training of Deep Convolutional Networks
Darryl D. Lin, Sachin S. Talathi, and V. Sreekanth Annapureddy · 2016
Later among the works it cites.
XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks
Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi · 2016
Later among the works it cites.
Training Neural Networks Without Gradients: A Scalable ADMM Approach
Gavin Taylor, Ryan Burmeister, Zheng Xu, Bharat Singh, Ankit Patel, and Tom Goldstein · 2016
Later among the works it cites.
DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Recursive Decomposition for Nonconvex Optimization
Abram L. Friesen and Pedro Domingos · 2015
Cited alongside, same era.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Lei Ba · 2015
Cited alongside, same era.
Difference target propagation
Dong Hyun Lee, Saizheng Zhang, Asja Fischer, and Yoshua Bengio · 2015
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Fei-Fei Li · 2015
Cited alongside, same era.
The Sum-Product Theorem: A Foundation for Learning Tractable Models
Abram L. Friesen and Pedro Domingos · 2016
Cited alongside, same era.
Binarized Neural Networks
Itay Hubara, Daniel Soudry, and Ran El-Yaniv · 2016
Cited alongside, same era.
Shuchang Zhou, Yuxin Wu, Zekun Ni, Xinyu Zhou, He Wen, and Yuheng Zou · 2016
Later among the works it cites.
Training Quantized Nets: A Deeper Understanding
Hao Li, Soham De, Zheng Xu, Christoph Studer, Hanan Samet, and Tom Goldstein · 2017
Closest in time.
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaev, Ganesh Venkatesh, and Hao Wu · 2017
Closest in time.
Failures of Gradient-Based Deep Learning
Shai Shalev-Shwartz, Ohad Shamir, and Shaked Shammah · 2017
Closest in time.
How to Train a Compact Binary Neural Network with High Accuracy ?
Wei Tang, Gang Hua, and Liang Wang · 2017
Closest in time.
Trained Ternary Quantization
Chenzhuo Zhu, Song Han, Huizi Mao, and William J. Dally · 2017
Closest in time.