Fetching the paper…
Reading the bibliography…
We consider optimization problems in which the objective requires an inner loop with many steps or is the limit of a sequence of increasingly costly approximations.
Matrix inversion by a Monte Carlo method
Forsythe, G. E. and Leibler, R. A · 1950
Earlier work this paper cites.
Use of different Monte Carlo sampling techniques
Kahn, H · 1955
Earlier work this paper cites.
Monte Carlo Principles and Neutron Transport Problems
Spanier, J. and Gelbard, E. M · 1969
Earlier work this paper cites.
Principles of Mathematical Analysis , volume 3
Rudin, W. et al · 1976
Earlier work this paper cites.
Stochastic method for the numerical study of lattice fermions
Kuti, J · 1982
Earlier work this paper cites.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
Schmidhuber, J · 1987
Earlier work this paper cites.
Unbiased Monte Carlo evaluation of certain functional integrals
Wagner, W · 1987
Earlier work this paper cites.
Aerodynamic design via control theory
Jameson, A · 1988
Earlier work this paper cites.
Particle transport and image synthesis
Arvo, J. and Kirk, D · 1990
Earlier work this paper cites.
Unbiased nonparametric estimation of the derivative of the mean
Rychlik, T · 1990
Earlier work this paper cites.
On the optimization of a synaptic learning rule
Bengio, S., Bengio, Y., Cloutier, J., and Gecsei, J · 1992
Earlier work this paper cites.
A class of unbiased kernel estimates of a probability density function
Rychlik, T · 1995
Earlier work this paper cites.
Reverse accumulation and implicit functions
Christianson, B · 1998
Earlier work this paper cites.
Particle filters for partially observed diffusions
Fearnhead, P., Papaspiliopoulos, O., and Roberts, G. O · 2008
Earlier work this paper cites.
A general method for debiasing a Monte Carlo estimator
McLeish, D · 2010
Earlier work this paper cites.
A new approach to unbiased estimation for SDEs
Rhee, C.-h. and Glynn, P. W · 2012
Cited alongside, same era.
Playing Russian roulette with intractable likelihoods
Girolami, M., Lyne, A.-M., Strathmann, H., Simpson, D., and Atchade, Y · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Cited alongside, same era.
Auto-encoding variational Bayes
Kingma, D. P. and Welling, M · 2014
Cited alongside, same era.
On Russian roulette estimates for Bayesian inference with doubly-intractable likelihoods
Lyne, A.-M., Girolami, M., Atchadé, Y., Strathmann, H., Simpson, D., et al · 2015
Cited alongside, same era.
Gradient-based hyperparameter optimization through reversible learning
Maclaurin, D., Duvenaud, D., and Adams, R · 2015
Cited alongside, same era.
Regularizing and optimizing LSTM language models
Merity, S., Keskar, N. S., and Socher, R · 2017
Later among the works it cites.
Unbiasing truncated backpropagation through time
Tallec, C. and Ollivier, Y · 2017
Later among the works it cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Later among the works it cites.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J · 2018
Later among the works it cites.
SMASH: One-shot model architecture search through hypernetworks
Brock, A., Lim, T., Ritchie, J., and Weston, N · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Unbiased estimation with square root convergence for SDE models
Rhee, C.-h. and Glynn, P. W · 2015
Cited alongside, same era.
Coupling adaptive batch sizes with learning rates
Balles, L., Romero, J., and Hennig, P · 2016
Cited alongside, same era.
Introduction to online convex optimization
Hazan, E. et al · 2016
Cited alongside, same era.
Optimization as a model for few-shot learning
Ravi, S. and Larochelle, H · 2016
Cited alongside, same era.
Markov chain truncation for doubly-intractable inference
Wei, C. and Murray, I · 2016
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Cremer, C., Li, X., and Duvenaud, D · 2018
Later among the works it cites.
Semi-amortized variational autoencoders
Kim, Y., Wiseman, S., Miller, A. C., Sontag, D., and Rush, A. M · 2018
Later among the works it cites.
Stochastic hyperparameter optimization through hypernetworks
Lorraine, J. and Duvenaud, D · 2018
Later among the works it cites.
An analysis of neural language modeling at multiple scales
Merity, S., Keskar, N. S., and Socher, R · 2018
Later among the works it cites.
Learned optimizers that outperform SGD on wall-clock and validation loss
Metz, L., Maheswaranathan, N., Nixon, J., Freeman, C. D., and Sohl-Dickstein, J · 2018
Later among the works it cites.
Truncated back-propagation for bilevel optimization
Shaban, A., Cheng, C.-A., Hatch, N., and Boots, B · 2018
Later among the works it cites.
Learning longer-term dependencies in RNNs with auxiliary losses
Trinh, T. H., Dai, A. M., Luong, T., and Le, Q. V · 2018
Later among the works it cites.
Understanding short-horizon bias in stochastic meta-optimization
Wu, Y., Ren, M., Liao, R., and Grosse., R · 2018
Later among the works it cites.
Credit assignment techniques in stochastic computation graphs
Weber, T., Heess, N., Buesing, L., and Silver, D · 2019
Closest in time.