Fetching the paper…
Reading the bibliography…
Backpropagation, the cornerstone of deep learning, is limited to computing gradients for continuous variables.
Statistical theory of extreme values and some practical applications : A series of lectures
Gumbel, E. J · 1954
Earlier work this paper cites.
The perceptron, a perceiving and recognizing automaton Project Para
Rosenblatt, F · 1957
Earlier work this paper cites.
Principles of neurodynamics
Mullin, A. A. and Rosenblatt, F · 1962
Earlier work this paper cites.
Learning representations by backpropagating errors
Rumelhari, D. E., Hintont, G. E., Ronald, J., and Williams · 1986
Earlier work this paper cites.
Connectionist learning of belief networks
Neal, R. M · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
The helmholtz machine
Dayan, P., Hinton, G. E., Neal, R. M., and Zemel, R. S · 1995
Earlier work this paper cites.
Computer methods for ordinary differential equations and differential-algebraic equations
Ascher, U. M. and Petzold, L. R · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
The optimal reward baseline for gradient-based reinforcement learning
Weaver, L. and Tao, N · 2001
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A. C · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Deep autoregressive networks
Gregor, K., Danihelka, I., Mnih, A., Blundell, C., and Wierstra, D · 2014
Earlier work this paper cites.
A* sampling
Maddison, C. J., Tarlow, D., and Minka, T. P · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
Human-level concept learning through probabilistic program induction
Lake, B. M., Salakhutdinov, R., and Tenenbaum, J. B · 2015
Cited alongside, same era.
Muprop: Unbiased backpropagation for stochastic neural networks
Gu, S. S., Levine, S., Sutskever, I., and Mnih, A · 2016
Cited alongside, same era.
Self-critical sequence training for image captioning
Rennie, S. J., Marcheret, E., Mroueh, Y., Ross, J., and Goel, V · 2016
Cited alongside, same era.
Learning to compose task-specific tree structures
Choi, J., Yoo, K. M., and goo Lee, S · 2017
Cited alongside, same era.
Backpropagation through the void: Optimizing control variates for black-box gradient estimation
Grathwohl, W., Choi, D., Wu, Y., Roeder, G., and Duvenaud, D. K · 2018
Later among the works it cites.
Listops: A diagnostic dataset for latent tree learning
Nangia, N. and Bowman, S. R · 2018
Later among the works it cites.
Searching for a robust neural architecture in four gpu hours
Dong, X. and Yang, Y · 2019
Later among the works it cites.
Darts: Differentiable architecture search
Liu, H., Simonyan, K., and Yang, Y · 2019
Later among the works it cites.
On the variance of the adaptive learning rate and beyond
Liu, L., Jiang, H., He, P., Chen, W., Liu, X., Gao, J., and Han, J · 2020
Later among the works it cites.
Low bias low variance gradient estimates for boolean stochastic networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hierarchical multiscale recurrent neural networks
Chung, J., Ahn, S., and Bengio, Y · 2017
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Jang, E., Gu, S. S., and Poole, B · 2017
Cited alongside, same era.
Evaluating the variance of likelihood-ratio gradient estimators
Tokui, S. and Sato, I · 2017
Cited alongside, same era.
Rebar: Low-variance, unbiased gradient estimates for discrete latent variable models
Tucker, G., Mnih, A., Maddison, C. J., Lawson, J., and Sohl-Dickstein, J. N · 2017
Cited alongside, same era.
Neural discrete representation learning
van den Oord, A., Vinyals, O., and Kavukcuoglu, K · 2017
Cited alongside, same era.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Xiao, H., Rasul, K., and Vollgraf, R · 2017
Cited alongside, same era.
Pervez, A., Cohen, T., and Gavves, E · 2020
Later among the works it cites.
Coupled gradient estimators for discrete latent variables
Dong, Z., Mnih, A., and Tucker, G · 2021
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N. M · 2021
Later among the works it cites.
Rao-blackwellizing the straight-through gumbel-softmax gradient estimator
Paulus, M. B., Maddison, C. J., and Krause, A · 2021
Later among the works it cites.
Training discrete deep generative models via gapped straight-through estimator
Fan, T.-H., Chi, T.-C., Rudnicky, A. I., and Ramadge, P. J · 2022
Later among the works it cites.
Gradient estimation with discrete stein operators
Shi, J., Zhou, Y., Hwang, J., Titsias, M., and Mackey, L · 2022
Later among the works it cites.
Sparse backpropagation for moe training
Liu, L., Gao, J., and Chen, W · 2023
Closest in time.