Fetching the paper…
Reading the bibliography…
We present a scalable and effective exploration strategy based on Thompson sampling for reinforcement learning (RL).
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
Frequentist regret bounds for randomized least-squares value iteration
Andrea Zanette, David Brandfonbrener, Emma Brunskill, Matteo Pirotta, and Alessandro Lazaric · 1964
Earlier work this paper cites.
Handbook of mathematical functions with formulas, graphs, and mathematical tables , volume 55
Milton Abramowitz and Irene A Stegun · 1964
Earlier work this paper cites.
Brownian dynamics as smart monte carlo simulation
Peter J Rossky, Jimmie D Doll, and Harold L Friedman · 1978
Earlier work this paper cites.
Exponential convergence of langevin distributions and their discrete approximations
Gareth O Roberts and Richard L Tweedie · 1996
Earlier work this paper cites.
A bayesian framework for reinforcement learning
Malcolm Strens · 2000
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Langevin diffusions and metropolis-hastings algorithms
Gareth O Roberts and Osnat Stramer · 2002
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Alexander L Strehl and Michael L Littman · 2008
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Double q-learning
Hado Van Hasselt · 2010
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
Mcmc using hamiltonian dynamics
Radford M Neal et al · 2011
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman, Geoffrey Hinton, et al · 2012
Earlier work this paper cites.
Matrix analysis
Roger A Horn and Charles R Johnson · 2012
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Thompson sampling for contextual bandits with linear payoffs
Shipra Agrawal and Navin Goyal · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Analysis and geometry of Markov diffusion operators , volume 103
Dominique Bakry, Ivan Gentil, Michel Ledoux, et al · 2014
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
On the convergence of stochastic gradient mcmc algorithms with high-order integrators
Changyou Chen, Nan Ding, and Lawrence Carin · 2015
Earlier work this paper cites.
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Earlier work this paper cites.
Preconditioned stochastic gradient langevin dynamics for deep neural networks
Chunyuan Li, Changyou Chen, David Carlson, and Lawrence Carin · 2016
Earlier work this paper cites.
Variance reduction in stochastic gradient langevin dynamics
Kumar Avinava Dubey, Sashank J Reddi, Sinead A Williamson, Barnabas Poczos, Alexander J Smola, and Eric P Xing · 2016
Earlier work this paper cites.
Consistency and fluctuations for stochastic gradient langevin dynamics
Yee Whye Teh, Alexandre H Thiery, and Sebastian J Vollmer · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Earlier work this paper cites.
Why is posterior sampling better than optimism for reinforcement learning?
Ian Osband and Benjamin Van Roy · 2017
Cited alongside, same era.
Count-based exploration with neural density models
Georg Ostrovski, Marc G Bellemare, Aäron Oord, and Rémi Munos · 2017
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
# exploration: A study of count-based exploration for deep reinforcement learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, OpenAI Xi Chen, Yan Duan, John Schulman, Filip DeTurck, and Pieter Abbeel · 2017
Cited alongside, same era.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Shipra Agrawal and Randy Jia · 2017
Cited alongside, same era.
Linear thompson sampling revisited
Marc Abeille and Alessandro Lazaric · 2017
A finite-time analysis of two time-scale actor-critic methods
Yue Frank Wu, Weitong Zhang, Pan Xu, and Quanquan Gu · 2020
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvari, Mengdi Wang, and Lin Yang · 2020
Later among the works it cites.
Reinforcement learning with general value function approximation: Provably efficient approach via bounded eluder dimension
Ruosong Wang, Russ R Salakhutdinov, and Lin Yang · 2020
Later among the works it cites.
Learning with good feature representations in bandits and in rl with a generative model
Tor Lattimore, Csaba Szepesvari, and Gellert Weisz · 2020
Later among the works it cites.
On worst-case regret of linear thompson sampling
Nima Hamidi and Mohsen Bayati · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Theoretical guarantees for approximate sampling from smooth and log-concave densities
Arnak S Dalalyan · 2017
Cited alongside, same era.
Noisy networks for exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Ian Osband, Alex Graves, Vlad Mnih, Remi Munos, Demis Hassabis, Olivier Pietquin, et al · 2017
Cited alongside, same era.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Efficient exploration through bayesian deep q-networks
Kamyar Azizzadenesheli, Emma Brunskill, and Animashree Anandkumar · 2018
Cited alongside, same era.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Cited alongside, same era.
Samarth Sinha, Homanga Bharadhwaj, Aravind Srinivas, and Animesh Garg · 2020
Later among the works it cites.
Vikranth Dwaracherla and Benjamin Van Roy · 2020
Later among the works it cites.
DQN Zoo: Reference implementations of DQN-based agents, 2020
John Quan and Georg Ostrovski · 2020
Later among the works it cites.
Randomized exploration in reinforcement learning with general value function approximation
Haque Ishfaq, Qiwen Cui, Viet Nguyen, Alex Ayoub, Zhuoran Yang, Zhaoran Wang, Doina Precup, and Lin Yang · 2021
Later among the works it cites.
Nearly minimax optimal reinforcement learning for linear mixture markov decision processes
Dongruo Zhou, Quanquan Gu, and Csaba Szepesvari · 2021
Later among the works it cites.
Provably correct optimization and exploration with non-linear policies
Fei Feng, Wotao Yin, Alekh Agarwal, and Lin Yang · 2021
Later among the works it cites.
Faster convergence of stochastic gradient langevin dynamics for non-log-concave sampling
Difan Zou, Pan Xu, and Quanquan Gu · 2021
Later among the works it cites.
Training larger networks for deep reinforcement learning
Kei Ota, Devesh K Jha, and Asako Kanezaki · 2021
Later among the works it cites.
Revisiting rainbow: Promoting more insightful and inclusive deep reinforcement learning research
Johan Samir Obando Ceron and Pablo Samuel Castro · 2021
Later among the works it cites.
Hyperdqn: A randomized exploration method for deep reinforcement learning
Ziniu Li, Yingru Li, Yushun Zhang, Tong Zhang, and Zhi-Quan Luo · 2021
Later among the works it cites.
A provably efficient model-free posterior sampling method for episodic reinforcement learning
Christoph Dann, Mehryar Mohri, Tong Zhang, and Julian Zimmert · 2021
Later among the works it cites.
Langevin monte carlo for contextual bandits
Pan Xu, Hongkai Zheng, Eric V Mazumdar, Kamyar Azizzadenesheli, and Animashree Anandkumar · 2022
Later among the works it cites.
Stochastic gradient langevin dynamics with adaptive drifts
Sehwan Kim, Qifan Song, and Faming Liang · 2022
Later among the works it cites.
Cliff diving: Exploring reward surfaces in reinforcement learning environments
Ryan Sullivan, Justin K Terry, Benjamin Black, and John P Dickerson · 2022
Later among the works it cites.
Making linear mdps practical via contrastive representation learning
Tianjun Zhang, Tongzheng Ren, Mengjiao Yang, Joseph Gonzalez, Dale Schuurmans, and Bo Dai · 2022
Later among the works it cites.
Near-optimal randomized exploration for tabular markov decision processes
Zhihan Xiong, Ruoqi Shen, Qiwen Cui, Maryam Fazel, and Simon Shaolei Du · 2022
Later among the works it cites.
Feel-good thompson sampling for contextual bandits and reinforcement learning
Tong Zhang · 2022
Later among the works it cites.
From dirichlet to rubin: Optimistic exploration in rl without bonuses
Daniil Tiapkin, Denis Belomestny, Éric Moulines, Alexey Naumov, Sergey Samsonov, Yunhao Tang, Michal Valko, and Pierre Ménard · 2022
Later among the works it cites.
Tianshou: A highly modularized deep reinforcement learning library
Jiayi Weng, Huayu Chen, Dong Yan, Kaichao You, Alexis Duburcq, Minghao Zhang, Yi Su, Hang Su, and Jun Zhu · 2022
Later among the works it cites.
Nearly minimax optimal reinforcement learning for linear markov decision processes
Jiafan He, Heyang Zhao, Dongruo Zhou, and Quanquan Gu · 2023
Closest in time.
Toshinori Kitamura, Tadashi Kozuno, Yunhao Tang, Nino Vieillard, Michal Valko, Wenhao Yang, Jincheng Mei, Pierre Ménard, Mohammad Gheshlaghi Azar, Rémi Munos, et al · 2023
Closest in time.
Maximize to explore: One objective function fusing estimation, planning, and exploration
Zhihan Liu, Miao Lu, Wei Xiong, Han Zhong, Hao Hu, Shenao Zhang, Sirui Zheng, Zhuoran Yang, and Zhaoran Wang · 2023
Closest in time.
Bilinear exponential family of mdps: frequentist regret bound with tractable exploration & planning
Reda Ouhamma, Debabrota Basu, and Odalric Maillard · 2023
Closest in time.
Online RL in Linearly q π q^{\pi} -Realizable MDPs Is as Easy as in Linear MDPs If You Learn What to Ignore
Gellért Weisz, András György, and Csaba Szepesvari · 2023
Closest in time.
Posterior sampling with delayed feedback for reinforcement learning with linear function approximation
Nikki Lijing Kuang, Ming Yin, Mengdi Wang, Yu-Xiang Wang, and Yian Ma · 2024
Closest in time.