Fetching the paper…
Reading the bibliography…
We introduce the active exploration problem in Markov decision processes (MDPs).
Algorithmic complexity: three np-hard problems in computational statistics
Welch, W. J. (1982) · 1982
Earlier work this paper cites.
Geometric bounds for eigenvalues of markov chains
Diaconis, P., Stroock, D., et al. (1991) · 1991
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L. (1994) · 1994
Earlier work this paper cites.
Fastest mixing markov chain on a graph
Boyd, S., Diaconis, P., and Xiao, L. (2004) · 2004
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Chentanez, N., Barto, A. G., and Singh, S. P. (2005) · 2005
Earlier work this paper cites.
Optimal Design of Experiments
Pukelsheim, F. (2006) · 2006
Earlier work this paper cites.
Natural actor-critic algorithms
Bhatnagar, S., Sutton, R., Ghavamzadeh, M., and Lee, M. (2009) · 2009
Earlier work this paper cites.
Active learning in heteroscedastic noise
Antos, A., Grover, V., and Szepesvári, C. (2010) · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P. (2010) · 2010
Cited alongside, same era.
Models for autonomously motivated exploration in reinforcement learning
Auer, P., Lim, S. H., and Watkins, C. (2011) · 2011
Cited alongside, same era.
Upper-confidence-bound algorithms for active learning in multi-armed bandits
Carpentier, A., Lazaric, A., Ghavamzadeh, M., Munos, R., and Auer, P. (2011) · 2011
Cited alongside, same era.
Patrolling security games: Definition and algorithms for solving large instances with single patroller and single intruder
Basilico, N., Gatti, N., and Amigoni, F. (2012) · 2012
Cited alongside, same era.
The steady-state control problem for markov decision processes
Akshay, S., Bertrand, N., Haddad, S., and Helouet, L. (2013) · 2013
Cited alongside, same era.
Revisiting frank-wolfe: Projection-free sparse convex optimization
Mixing time estimation in reversible markov chains from a single sample path
Hsu, D. J., Kontorovich, A., and Szepesvári, C. (2015) · 2015
Later among the works it cites.
Concentration inequalities for markov chains by marton couplings and spectral methods
Paulin, D. et al. (2015) · 2015
Later among the works it cites.
Fast rates for bandit optimization with upper-confidence frank-wolfe
Berthet, Q. and Perchet, V. (2017) · 2017
Later among the works it cites.
Optimal policies for observing time series and related restless bandit problems
Dance, C. R. and Silander, T. (2017) · 2017
Later among the works it cites.
Provably efficient maximum entropy exploration
Hazan, E., Kakade, S. M., Singh, K., and Soest, A. V. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jaggi, M. (2013) · 2013
Cited alongside, same era.
Theory of disagreement-based active learning
Hanneke, S. (2014) · 2014
Cited alongside, same era.
Commitment without regrets: Online learning in stackelberg security games
Balcan, M.-F., Blum, A., Haghtalab, N., and Procaccia, A. D. (2015) · 2015
Cited alongside, same era.
Rolf, E., Fridovich-Keil, D., Simchowitz, M., Recht, B., and Tomlin, C. (2018) · 2018
Later among the works it cites.
Bandit Algorithms
Lattimore, T. and Szepesvári, C. (2019) · 2019
Closest in time.