Fetching the paper…
Reading the bibliography…
We study the problem of efficient exploration in order to learn an accurate model of an environment, modeled as a Markov decision process (MDP).
Wang Chi Cheung · 1905
Earlier work this paper cites.
Weighted entropy
Silviu Guiaşu · 1971
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Tuning bandit algorithms in stochastic environments
Jean-Yves Audibert, Rémi Munos, and Csaba Szepesvári · 2007
Earlier work this paper cites.
An analysis of model-based interval estimation for Markov decision processes
Alexander L Strehl and Michael L Littman · 2008
Earlier work this paper cites.
Natural actor–critic algorithms
Shalabh Bhatnagar, Richard S Sutton, Mohammad Ghavamzadeh, and Mark Lee · 2009
Earlier work this paper cites.
Empirical Bernstein bounds and sample variance penalization
Andreas Maurer and Massimiliano Pontil · 2009
Earlier work this paper cites.
Active learning of MDP models
Mauricio Araya-López, Olivier Buffet, Vincent Thomas, and François Charpillet · 2011
Cited alongside, same era.
Autonomous exploration for navigating in MDPs
Shiau Hong Lim and Peter Auer · 2012
Cited alongside, same era.
Revisiting Frank-Wolfe: Projection-free sparse convex optimization
Martin Jaggi · 2013
Cited alongside, same era.
Markov Decision Processes.: Discrete Stochastic Dynamic Programming
Martin L Puterman · 2014
Cited alongside, same era.
Concentration inequalities for Markov chains by Marton couplings and spectral methods
Daniel Paulin · 2015
Cited alongside, same era.
First-order methods in optimization , volume 25
Amir Beck · 2017
Cited alongside, same era.
Improved analysis of UCRL2B, 2019
Ronan Fruit, Matteo Pirotta, and Alessandro Lazaric · 2019
Later among the works it cites.
Provably efficient maximum entropy exploration
Elad Hazan, Sham Kakade, Karan Singh, and Abby Van Soest · 2019
Later among the works it cites.
Online convex optimization in adversarial Markov decision processes
Aviv Rosenberg and Yishay Mansour · 2019
Later among the works it cites.
Active exploration in Markov decision processes
Jean Tarbouriech and Alessandro Lazaric · 2019
Later among the works it cites.
Reward-free exploration for reinforcement learning
Chi Jin, Akshay Krishnamurthy, Max Simchowitz, and Tiancheng Yu · 2020
Closest in time.
Adaptive sampling for estimating probability distributions
Shubhanshu Shekhar, Mohammad Ghavamzadeh, and Tara Javidi · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fast rates for bandit optimization with upper-confidence Frank-Wolfe
Quentin Berthet and Vianney Perchet · 2017
Cited alongside, same era.
Regret minimization for reinforcement learning with vectorial feedback and complex objectives
Wang Chi Cheung
Cited in the paper.