Fetching the paper…
Reading the bibliography…
In this paper, we formalise order-robust optimisation as an instance of online learning minimising simple regret, and propose Vroom, a zero'th order optimisation algorithm capable of achieving vanishing regret in non-stationary environments, while recovering favorable rates under stochastic reward-generating processes.
Scalable and order-robust continual learning with hierarchically decomposed networks
Yoon, J., Kim, S., Yang, E., and Hwang, S. J. (2019) · 1902
Earlier work this paper cites.
Random optimization
Matyas, J. (1965) · 1965
Earlier work this paper cites.
On tail probabilities for martingales
Freedman, D. A. (1975) · 1975
Earlier work this paper cites.
Lifelong robot learning
Thrun, S. and Mitchell, T. M. (1995) · 1995
Earlier work this paper cites.
Global Optimization in Action. Continous and Lipschitz Optimization: Algorithms, Implementations and Applications
Pintér, J. D. (1996) · 1996
Earlier work this paper cites.
Applications and explanations of Zipf’s law
Powers, D. (1998) · 1998
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
French, R. M. (1999) · 1999
Earlier work this paper cites.
Global Optimization with Non-Convex Constraints: Sequential and Parallel Algorithms
Strongin, R. and Sergeyev, Y. (2000) · 2000
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Auer, P. (2002) · 2002
Earlier work this paper cites.
Global optimization using interval analysis: revised and expanded
Hansen, E. and Walster, G. W. (2003) · 2003
Earlier work this paper cites.
Bandit-based Monte-Carlo planning
Kocsis, L. and Szepesvári, C. (2006) · 2006
Earlier work this paper cites.
Improved rates for the stochastic continuum-armed bandit problem
Auer, P., Ortner, R., and Szepesvári, C. (2007) · 2007
Earlier work this paper cites.
Bandit algorithms for tree search
Coquelin, P.-A. and Munos, R. (2007) · 2007
Earlier work this paper cites.
Inequalities on the lambert w function and hyperpower function
Hoorfar, A. and Hassani, M. (2008) · 2008
Earlier work this paper cites.
Optimistic Planning of Deterministic Systems
Hren, J.-F. and Munos, R. (2008) · 2008
Cited alongside, same era.
Multi-armed bandit problems in metric spaces
Kleinberg, R., Slivkins, A., and Upfal, E. (2008) · 2008
Cited alongside, same era.
Pure exploration in multi-armed bandit problems
Bubeck, S., Munos, R., and Stoltz, G. (2009) · 2009
Cited alongside, same era.
Empirical bernstein bounds and sample variance penalization
Maurer, A. and Pontil, M. (2009) · 2009
Cited alongside, same era.
Best arm identification in multi-armed bandits
Audibert, J.-Y., Bubeck, S., and Munos, R. (2010) · 2010
Cited alongside, same era.
Open-loop optimistic planning
Bubeck, S. and Munos, R. (2010) · 2010
Cited alongside, same era.
Online multi-task learning for policy gradient methods
Ammar, H. B., Eaton, E., Ruvolo, P., and Taylor, M. (2014) · 2014
Later among the works it cites.
Online stochastic optimization under correlated bandit feedback
Azar, M. G., Lazaric, A., and Brunskill, E. (2014) · 2014
Later among the works it cites.
From bandits to Monte-Carlo tree search: The optimistic principle applied to optimization and planning
Munos, R. (2014) · 2014
Later among the works it cites.
Introduction to online convex optimization
Hazan, E. et al. (2016) · 2016
Later among the works it cites.
Global continuous optimization with error bound and fast convergence
Kawaguchi, K., Maruyama, Y., and Zheng, X. (2016) · 2016
Later among the works it cites.
Kernel-based methods for bandit convex optimization
Bubeck, S., Lee, Y. T., and Eldan, R. (2017) · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zinkevich, M. (2003) · 2010
Cited alongside, same era.
X-armed bandits
Bubeck, S., Munos, R., Stoltz, G., and Szepesvári, C. (2011) · 2011
Cited alongside, same era.
Optimistic optimization of deterministic functions without the knowledge of its smoothness
Munos, R. (2011) · 2011
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Bubeck, S. and Cesa-Bianchi, N. (2012) · 2012
Cited alongside, same era.
Exponential regret bounds for Gaussian process bandits with deterministic observations
de Freitas, N., Smola, A., and Zoghi, M. (2012) · 2012
Cited alongside, same era.
Kullback–leibler upper confidence bounds for optimal sequential allocation
Cappé, O., Garivier, A., Maillard, O.-A., Munos, R., Stoltz, G., et al. (2013) · 2013
Cited alongside, same era.
Later among the works it cites.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al. (2017) · 2017
Later among the works it cites.
Random gradient-free minimization of convex functions
Nesterov, Y. and Spokoiny, V. (2017) · 2017
Later among the works it cites.
Best of both worlds: Stochastic & adversarial best-arm identification
Abbasi-Yadkori, Y., Bartlett, P., Gabillon, V., Malek, A., and Valko, M. (2018) · 2018
Later among the works it cites.
Adaptivity to Smoothness in X-armed bandits
Locatelli, A. and Carpentier, A. (2018) · 2018
Later among the works it cites.
A simple parameter-free and adaptive approach to optimization under a minimal local smoothness assumption
Bartlett, P. L., Gabillon, V., and Valko, M. (2019) · 2019
Closest in time.
Continual lifelong learning with neural networks: A review
Parisi, G. I., Kemker, R., Part, J. L., Kanan, C., and Wermter, S. (2019) · 2019
Closest in time.
General parallel optimization without metric
Shang, X., Kaufmann, E., and Valko, M. (2019) · 2019
Closest in time.