Fetching the paper…
Reading the bibliography…
We propose UCBMQ, Upper Confidence Bound Momentum Q-learning, a new algorithm for reinforcement learning in tabular and possibly stage-dependent, episodic Markov decision process.
Q-learning
Watkins, Chris J. and Dayan, Peter · 1992
Earlier work this paper cites.
Efficient reinforcement learning
Fiechter, Claude Nicolas · 1994
Earlier work this paper cites.
Markov Decision Processes - Discrete Stochastic Dynamic Programming
Puterman, Martin L · 1994
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, Richard S. and Barto, Andrew G · 1998
Earlier work this paper cites.
On the Sample Complexity of Reinforcement Learning
Kakade, Sham · 2003
Earlier work this paper cites.
Almost optimal model-free reinforcement learning via reference-advantage decomposition
Zhang, Zihan, Zhou, Yuan, and Ji, Xiangyang · 2004
Earlier work this paper cites.
Zhang, Zihan, Ji, Xiangyang, and Du, Simon S · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, Thomas, Ortner, Ronald, and Auer, Peter · 2010
Earlier work this paper cites.
Model-based reinforcement learning with nearly tight exploration complexity bounds
Szita, István and Szepesvári, Csaba · 2010
Cited alongside, same era.
Speedy Q-learning
Azar, Mohammad Gheshlaghi, Munos, Remi, Ghavamzadeh, Mohammad, and Kappen, Hilbert J · 2011
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, Mohammad Gheshlaghi, Osband, Ian, and Munos, Rémi · 2017
Cited alongside, same era.
Unifying PAC and regret: Uniform PAC bounds for episodic reinforcement learning
Dann, Christoph, Lattimore, Tor, and Brunskill, Emma · 2017
Cited alongside, same era.
Efficient bias-span-constrained exploration-exploitation in reinforcement learning
Fruit, Ronan, Pirotta, Matteo, Lazaric, Alessandro, and Ortner, Ronald · 2018
Cited alongside, same era.
Zanette, Andrea and Brunskill, Emma · 2019
Later among the works it cites.
Regret bounds for kernel-based reinforcement learning
Domingues, Omar D., Ménard, Pierre, Pirotta, Matteo, Kaufmann, Emilie, and Valko, Michal · 2020
Later among the works it cites.
Adaptive discretization for model-based reinforcement learning
Sinclair, Sean R., Wang, Tianyu, Jain, Gauri, Banerjee, Siddhartha, and Yu, Christina Lee · 2020
Later among the works it cites.
Is long horizon reinforcement learning more difficult than short horizon reinforcement learning?
Wang, Ruosong, Du, Simon S., Yang, Lin F., and Kakade, Sham M · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Is Q-learning provably efficient?
Jin, Chi, Allen-Zhu, Zeyuan, Bubeck, Sebastien, and Jordan, Michael I · 2018
Cited alongside, same era.
Variance-aware regret bounds for undiscounted reinforcement learning in MDPs
Talebi, Mohammad Sadegh and Maillard, Odalric Ambrym · 2018
Cited alongside, same era.
Tight regret bounds for model-based reinforcement learning with greedy policies
Efroni, Yonathan, Merlis, Nadav, Ghavamzadeh, Mohammad, and Mannor, Shie · 2019
Cited alongside, same era.
rlberry - A Reinforcement Learning Library for Research and Education
Domingues, Omar Darwiche, Flet-Berliac, Yannis, Leurent, Edouard, Ménard, Pierre, Shang, Xuedong, and Valko, Michal
Cited in the paper.
Episodic reinforcement learning in finite MDPs: Minimax lower bounds revisited
Domingues, Omar Darwiche, Ménard, Pierre, Kaufmann, Emilie, and Valko, Michal
Cited in the paper.
Sidford, Aaron, Wang, Mengdi, Wu, Xian, Yang, Lin F., and Ye, Yinyu
Cited in the paper.
Variance reduced value iteration and faster algorithms for solving Markov decision processes
Sidford, Aaron, Wang, Mengdi, Wu, Xian, and Ye, Yinyu
Cited in the paper.
Weng, Bowen, Xiong, Huaqing, Zhao, Lin, Liang, Yingbin, and Zhang, Wei · 2020
Later among the works it cites.
Adaptive reward-free exploration
Kaufmann, Emilie, Ménard, Pierre, Domingues, Omar Darwiche, Jonsson, Anders, Leurent, Edouard, and Valko, Michal · 2021
Closest in time.
Fast active learning for pure exploration in reinforcement learning
Ménard, Pierre, Domingues, Omar Darwiche, Jonsson, Anders, Kaufmann, Emilie, Leurent, Edouard, and Valko, Michal · 2021
Closest in time.