Fetching the paper…
Reading the bibliography…
Neural networks with binary weights are computation-efficient and hardware-friendly, but their training is challenging because it involves a discrete optimization problem.
Mirror descent view for neural network quantization
Ajanthan, T., Gupta, K., Torr, P. H., Hartley, R., and Dokania, P. K · 1910
Earlier work this paper cites.
Evolutionsstrategie. optimierung technischer systeme nach prinzipien der biologischen evolution, 1976
Huning, A · 1976
Earlier work this paper cites.
Stochastic relaxation, Gibbs distributions, and the Bayesian restoration of images
Geman, S. and Geman, D · 1984
Earlier work this paper cites.
Exponential smoothing: The state of the art
Gardner Jr, E. S · 1985
Earlier work this paper cites.
Optimal information processing and bayes’s theorem
Zellner, A · 1988
Earlier work this paper cites.
Nondifferentiable optimization
Lemaréchal, C · 1989
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S.-I · 1998
Earlier work this paper cites.
An introduction to variational methods for graphical models
Jordan, M. I., Ghahramani, Z., Jaakkola, T. S., and Saul, L. K · 1999
Earlier work this paper cites.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Sparse Gaussian processes using pseudo-inputs
Snelson, E. and Ghahramani, Z · 2005
Earlier work this paper cites.
Graphical models, exponential families, and variational inference
Wainwright, M. J. and Jordan, M. I · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
MNIST handwritten digit database
LeCun, Y. and Cortes, C · 2010
Earlier work this paper cites.
Staines, J. and Barber, D · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A · 2013
Cited alongside, same era.
An empirical investigation of catastrophic forgetting in gradient-based neural networks
Goodfellow, I. J., Mirza, M., Xiao, D., Courville, A., and Bengio, Y · 2013
Cited alongside, same era.
Stochastic variational inference
Hoffman, M. D., Blei, D. M., Wang, C., and Paisley, J · 2013
Cited alongside, same era.
Spatially-sparse convolutional neural networks
Graham, B · 2014
Cited alongside, same era.
Binaryconnect: Training deep neural networks with binary weights during propagations
Courbariaux, M., Bengio, Y., and David, J.-P · 2015
Cited alongside, same era.
The concrete distribution: A continuous relaxation of discrete random variables
Maddison, C. J., Mnih, A., and Teh, Y. W · 2017
Later among the works it cites.
Continual learning through synaptic intelligence
Zenke, F., Poole, B., and Ganguli, S · 2017
Later among the works it cites.
Fast and scalable Bayesian deep learning by weight-perturbation in Adam
Khan, M. E., Nielsen, D., Tangkaratt, V., Lin, W., Gal, Y., and Srivastava, A · 2018
Later among the works it cites.
WRPN: wide reduced-precision networks
Mishra, A., Nurvitadhi, E., Cook, J. J., and Marr, D · 2018
Later among the works it cites.
Variational continual learning
Nguyen, C. V., Li, Y., Bui, T. D., and Turner, R. E · 2018
Later among the works it cites.
Probabilistic binary neural networks
Peters, J. W. and Welling, M · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
The information geometry of mirror descent
Raskutti, G. and Mukherjee, S · 2015
Cited alongside, same era.
A general framework for updating belief distributions
Bissiri, P. G., Holmes, C. C., and Walker, S. G · 2016
Cited alongside, same era.
Courbariaux, M., Hubara, I., Soudry, D., El-Yaniv, R., and Bengio, Y · 2016
Cited alongside, same era.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J. N., Pascanu, R., Rabinowitz, N. C., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., Hassabis, D., Clopath, C., Kumaran, D., and Hadsell, R · 2016
Cited alongside, same era.
Xnor-net: Imagenet classification using binary convolutional neural networks
Rastegari, M., Ordonez, V., Redmon, J., and Farhadi, A · 2016
Cited alongside, same era.
Later among the works it cites.
Learning discrete weights using the local reparameterization trick
Shayer, O., Levi, D., and Fetaya, E · 2018
Later among the works it cites.
An empirical study of binary neural networks’ optimisation
Alizadeh, M., Fernández-Marqués, J., Lane, N. D., and Gal, Y · 2019
Later among the works it cites.
Latent weights do not exist: Rethinking binarized neural network optimization
Helwegen, K., Widdicombe, J., Geiger, L., Liu, Z., Cheng, K.-T., and Nusselder, R · 2019
Later among the works it cites.
Relaxed quantization for discretized neural networks
Louizos, C., Reisser, M., Blankevoort, T., Gavves, E., and Welling, M · 2019
Later among the works it cites.
Practical deep learning with Bayesian principles
Osawa, K., Swaroop, S., Khan, M. E. E., Jain, A., Eschenhagen, R., Turner, R. E., and Yokota, R · 2019
Later among the works it cites.
Understanding straight-through estimator in training activation quantized neural nets
Yin, P., Lyu, J., Zhang, S., Osher, S., Qi, Y., and Xin, J · 2019
Later among the works it cites.
Meliusnet: Can binary neural networks achieve mobilenet-level accuracy?
Bethge, J., Bartz, C., Yang, H., Chen, Y., and Meinel, C · 2020
Closest in time.
Learning-Algorithms from Bayesian principles
Khan, M. E. and Rue, H · 2020
Closest in time.