Fetching the paper…
Reading the bibliography…
We present the Variational Adaptive Newton (VAN) method which is a black-box optimization method especially suitable for explorative-learning tasks such as active learning and reinforcement learning.
Stochastic relaxation, Gibbs distributions, and the Bayesian restoration of images
Geman, S. and Geman, D. (1984) · 1984
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S.-I. (1998) · 1998
Earlier work this paper cites.
Exploration and inference in learning from reinforcement
Wyatt, J. (1998) · 1998
Earlier work this paper cites.
Nonlinear programming
Bertsekas, D. P. (1999) · 1999
Earlier work this paper cites.
Adaptive stochastic approximation by the simultaneous perturbation method
Spall, J. C. (2000) · 2000
Earlier work this paper cites.
A Bayesian framework for reinforcement learning
Strens, M. (2000) · 2000
Earlier work this paper cites.
Completely derandomized self-adaptation in evolution strategies
Hansen, N. and Ostermeier, A. (2001) · 2001
Earlier work this paper cites.
Pattern recognition
Bishop, C. M. (2006) · 2006
Earlier work this paper cites.
Adaptive Newton-based multivariate smoothed functional algorithms for simulation optimization
Bhatnagar, S. (2007) · 2007
Earlier work this paper cites.
Active learning for logistic regression: an evaluation
Schein, A. I. and Ungar, L. H. (2007) · 2007
Earlier work this paper cites.
Graphical models, exponential families, and variational inference
Wainwright, M. J. and Jordan, M. I. (2008) · 2008
Earlier work this paper cites.
Natural evolution strategies
Wierstra, D., Schaul, T., Peters, J., and Schmidhuber, J. (2008) · 2008
Earlier work this paper cites.
Adaptive regularization of weight vectors
Crammer, K., Kulesza, A., and Dredze, M. (2009) · 2009
Earlier work this paper cites.
The Variational Gaussian Approximation Revisited
Opper, M. and Archambeau, C. (2009) · 2009
Earlier work this paper cites.
Brochu, E., Cora, V. M., and De Freitas, N. (2010) · 2010
Cited alongside, same era.
Exploring parameter space in reinforcement learning
Rückstieß, T., Sehnke, F., Schaul, T., Wierstra, D., Sun, Y., and Schmidhuber, J. (2010) · 2010
Cited alongside, same era.
LIBSVM: A library for support vector machines
Chang, C.-C. and Lin, C.-J. (2011) · 2011
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y. (2011) · 2011
Cited alongside, same era.
Bayesian Active Learning for Classification and Preference Learning
Houlsby, N., Huszár, F., Ghahramani, Z., and Lengyel, M. (2011) · 2011
Cited alongside, same era.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. A. (2014) · 2014
Later among the works it cites.
A globally convergent incremental Newton method
Gürbüzbalaban, M., Ozdaglar, A., and Parrilo, P. (2015) · 2015
Later among the works it cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2015) · 2015
Later among the works it cites.
The information geometry of mirror descent
Raskutti, G. and Mukherjee, S. (2015) · 2015
Later among the works it cites.
Variational inference with normalizing flows
Rezende, D. J. and Mohamed, S. (2015) · 2015
Later among the works it cites.
Openai gym
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Marlin, B., Khan, M., and Murphy, K. (2011) · 2011
Cited alongside, same era.
A tutorial introduction to Bayesian models of cognitive development
Perfors, A., Tenenbaum, J. B., Griffiths, T. L., and Xu, F. (2011) · 2011
Cited alongside, same era.
Bayesian exploration for intelligent identification of textures
Fishel, J. A. and Loeb, G. E. (2012) · 2012
Cited alongside, same era.
Variational Optimization
Staines, J. and Barber, D. (2012) · 2012
Cited alongside, same era.
Lecture 6.5-RMSprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G. (2012) · 2012
Cited alongside, same era.
Auto-encoding variational Bayes
Kingma, D. P. and Welling, M. (2013) · 2013
Cited alongside, same era.
Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
Ghadimi, S., Lan, G., and Zhang, H. (2014) · 2014
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Later among the works it cites.
A stochastic quasi-Newton method for large-scale optimization
Byrd, R. H., Hansen, S. L., Nocedal, J., and Singer, Y. (2016) · 2016
Later among the works it cites.
Adaptive Newton method for empirical risk minimization to statistical accuracy
Mokhtari, A., Daneshmand, H., Lucchi, A., Hofmann, T., and Ribeiro, A. (2016) · 2016
Later among the works it cites.
Noisy networks for exploration
Fortunato, M., Azar, M. G., Piot, B., Menick, J., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., Pietquin, O., et al. (2017) · 2017
Closest in time.
Deep Bayesian Active Learning with Image Data
Gal, Y., Islam, R., and Ghahramani, Z. (2017) · 2017
Closest in time.
Evolution Strategies, Variational Optimisation and Natural ES
Huszar, F. (2017) · 2017
Closest in time.
Khan, M. E. and Lin, W. (2017) · 2017
Closest in time.
Parameter space noise for exploration
Plappert, M., Houthooft, R., Dhariwal, P., Sidor, S., Chen, R. Y., Chen, X., Asfour, T., Abbeel, P., and Andrychowicz, M. (2017) · 2017
Closest in time.
Evolution Strategies as a Scalable Alternative to Reinforcement Learning
Salimans, T., Ho, J., Chen, X., Sidor, S., and Sutskever, I. (2017) · 2017
Closest in time.