Fetching the paper…
Reading the bibliography…
In this work, we propose KeRNS: an algorithm for episodic reinforcement learning in non-stationary Markov Decision Processes (MDPs) whose state-action set is endowed with a metric.
Efficient model-free reinforcement learning in metric spaces
Song, Z. and Sun, W. (2019) · 1905
Earlier work this paper cites.
Corruption robust exploration in episodic reinforcement learning
Lykouris, T., Simchowitz, M., Slivkins, A., and Sun, W. (2019) · 1911
Earlier work this paper cites.
Learning to act using real-time dynamic programming
Barto, A. G., Bradtke, S. J., and Singh, S. P. (1995) · 1995
Earlier work this paper cites.
Hidden-mode markov decision processes for nonstationary sequential decision making
Choi, S. P., Yeung, D.-Y., and Zhang, N. L. (2000) · 2000
Earlier work this paper cites.
Kernel-based reinforcement learning
Ormoneit, D. and Sen, Ś. (2002) · 2002
Earlier work this paper cites.
ε \varepsilon -mdps: Learning in varying environments
Szita, I., Takács, B., and Lörincz, A. (2002) · 2002
Earlier work this paper cites.
Regret Bounds for Kernel-Based Reinforcement Learning
Domingues, O. D., Ménard, P., Pirotta, M., Kaufmann, E., and Valko, M. (2020) · 2004
Earlier work this paper cites.
Discounted UCB
Kocsis, L. and Szepesvári, C. (2006) · 2006
Earlier work this paper cites.
Value function based reinforcement learning in changing markovian environments
Csáji, B. C. and Monostori, L. (2008) · 2008
Earlier work this paper cites.
Online markov decision processes
Even-Dar, E., Kakade, S. M., and Mansour, Y. (2009) · 2009
Earlier work this paper cites.
Online learning in markov decision processes with arbitrarily changing rewards and transitions
Yu, J. Y. and Mannor, S. (2009) · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P. (2010) · 2010
Earlier work this paper cites.
X-armed bandits
Bubeck, S., Munos, R., Stoltz, G., and Szepesvári, C. (2011) · 2011
Earlier work this paper cites.
On Upper-Confidence Bound Policies For Switching Bandit Problems
Garivier, A. and Moulines, E. (2011) · 2011
Cited alongside, same era.
Kernel-based reinforcement learning on representative states
Kveton, B. and Theocharous, G. (2012) · 2012
Cited alongside, same era.
Online regret bounds for undiscounted continuous reinforcement learning
Ortner, R. and Ryabko, D. (2012) · 2012
Cited alongside, same era.
Online markov decision processes under bandit feedback
Neu, G., György, A., Szepesvari, C., and Antos, A. (2013) · 2013
Cited alongside, same era.
Stochastic multi-armed-bandit problem with non-stationary rewards
Besbes, O., Gur, Y., and Zeevi, A. (2014) · 2014
Cited alongside, same era.
Online learning in markov decision processes with changing cost sequences
Dick, T., Gyorgy, A., and Szepesvari, C. (2014) · 2014
A new algorithm for non-stationary contextual bandits: Efficient, optimal and parameter-free
Chen, Y., Lee, C.-W., Luo, H., and Wei, C.-Y. (2019) · 2019
Later among the works it cites.
Tight regret bounds for model-based reinforcement learning with greedy policies
Efroni, Y., Merlis, N., Ghavamzadeh, M., and Mannor, S. (2019) · 2019
Later among the works it cites.
Bandits and experts in metric spaces
Kleinberg, R., Slivkins, A., and Upfal, E. (2019) · 2019
Later among the works it cites.
Non-stationary markov decision processes, a worst-case approach using model-based reinforcement learning
Lecarpentier, E. and Rachelson, E. (2019) · 2019
Later among the works it cites.
Online learning for markov decision processes in nonstationary environments: A dynamic regret analysis
Li, Y. and Li, N. (2019) · 2019
Later among the works it cites.
Variational regret bounds for reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L. (2014) · 2014
Cited alongside, same era.
Practical kernel-based reinforcement learning
Barreto, A. M., Precup, D., and Pineau, J. (2016) · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R. (2017) · 2017
Cited alongside, same era.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Dann, C., Lattimore, T., and Brunskill, E. (2017) · 2017
Cited alongside, same era.
Gajane, P., Ortner, R., and Auer, P. (2018) · 2018
Cited alongside, same era.
Is Q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I. (2018) · 2018
Cited alongside, same era.
Ortner, R., Gajane, P., and Auer, P. (2019) · 2019
Later among the works it cites.
Weighted linear bandits for non-stationary environments
Russac, Y., Vernade, C., and Cappé, O. (2019) · 2019
Later among the works it cites.
Adaptive discretization for episodic reinforcement learning in metric spaces
Sinclair, S. R., Banerjee, S., and Yu, C. L. (2019) · 2019
Later among the works it cites.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Zanette, A. and Brunskill, E. (2019) · 2019
Later among the works it cites.
Never give up: Learning directed exploration strategies
Badia, A. P., Sprechmann, P., Vitvitskyi, A., Guo, D., Piot, B., Kapturowski, S., Tieleman, O., Arjovsky, M., Pritzel, A., Bolt, A., and Blundell, C. (2020) · 2020
Closest in time.
Reinforcement learning for non-stationary Markov decision processes: The blessing of (More) optimism
Cheung, W. C., Simchi-Levi, D., and Zhu, R. (2020) · 2020
Closest in time.
Adaptive discretization for model-based reinforcement learning
Sinclair, S., Wang, T., Jain, G., Banerjee, S., and Yu, C. (2020) · 2020
Closest in time.