Fetching the paper…
Reading the bibliography…
For the problem of task-agnostic reinforcement learning (RL), an agent first collects samples from an unknown environment without the supervision of reward signals, then is revealed with a reward and is asked to compute a corresponding near-optimal policy.
Improvisation through physical understanding: Using novel objects as tools with visual foresight
Xie, A., Ebert, F., Levine, S., and Finn, C. (2019) · 1904
Earlier work this paper cites.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Simchowitz, M. and Jamieson, K. (2019) · 1905
Earlier work this paper cites.
Corruption robust exploration in episodic reinforcement learning
Lykouris, T., Simchowitz, M., Slivkins, A., and Sun, W. (2019) · 1911
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, R. I. and Tennenholtz, M. (2002) · 2002
Earlier work this paper cites.
Reward-free exploration for reinforcement learning
Jin, C., Krishnamurthy, A., Simchowitz, M., and Yu, T. (2020) · 2002
Earlier work this paper cites.
The sample complexity of exploration in the multi-armed bandit problem
Mannor, S. and Tsitsiklis, J. N. (2004) · 2004
Earlier work this paper cites.
Is long horizon reinforcement learning more difficult than short horizon reinforcement learning?
Wang, R., Du, S. S., Yang, L. F., and Kakade, S. M. (2020a) · 2005
Earlier work this paper cites.
Planning in markov decision processes with gap-dependent sample complexity
Jonsson, A., Kaufmann, E., Ménard, P., Domingues, O. D., Leurent, E., and Valko, M. (2020) · 2006
Earlier work this paper cites.
Adaptive reward-free exploration
Kaufmann, E., Ménard, P., Domingues, O. D., Jonsson, A., Leurent, E., and Valko, M. (2020) · 2006
Earlier work this paper cites.
On reward-free reinforcement learning with linear function approximation
Wang, R., Du, S. S., Yang, L. F., and Salakhutdinov, R. (2020b) · 2006
Earlier work this paper cites.
q q -learning with logarithmic regret
Yang, K., Yang, L. F., and Du, S. S. (2020) · 2006
Earlier work this paper cites.
Task-agnostic exploration in reinforcement learning
Zhang, X., Singla, A., et al. (2020a) · 2006
Earlier work this paper cites.
Fast active learning for pure exploration in reinforcement learning
Ménard, P., Domingues, O. D., Jonsson, A., Kaufmann, E., Leurent, E., and Valko, M. (2020) · 2007
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
Ortner, P. and Auer, R. (2007) · 2007
Earlier work this paper cites.
Optimistic linear programming gives logarithmic regret for irreducible mdps
Tewari, A. and Bartlett, P. L. (2007) · 2007
Cited alongside, same era.
Lenient regret for multi-armed bandits
Merlis, N. and Mannor, S. (2020) · 2008
Cited alongside, same era.
Empirical bernstein bounds and sample variance penalization
Maurer, A. and Pontil, M. (2009) · 2009
Cited alongside, same era.
Zhang, Z., Ji, X., and Du, S. S. (2020c) · 2009
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P. (2010) · 2010
Cited alongside, same era.
Provably efficient maximum entropy exploration
Hazan, E., Kakade, S. M., Singh, K., and Van Soest, A. (2018) · 2018
Later among the works it cites.
Open problem: The dependence of sample complexity lower bounds on planning horizon
Jiang, N. and Agarwal, A. (2018) · 2018
Later among the works it cites.
Is q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I. (2018) · 2018
Later among the works it cites.
Exploration in structured reinforcement learning
Ok, J., Proutiere, A., and Tranos, D. (2018) · 2018
Later among the works it cites.
Learning by playing solving sparse reward tasks from scratch
Riedmiller, M., Hafner, R., Lampe, T., Neunert, M., Degrave, J., Wiele, T., Mnih, V., Heess, N., and Springenberg, J. T. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nearly minimax optimal reward-free reinforcement learning
Zhang, Z., Du, S. S., and Ji, X. (2020b) · 2010
Cited alongside, same era.
Logarithmic regret for reinforcement learning with linear function approximation
He, J., Zhou, D., and Gu, Q. (2020) · 2011
Cited alongside, same era.
Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
Azar, M. G., Munos, R., and Kappen, H. J. (2013) · 2013
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
Dann, C. and Brunskill, E. (2015) · 2015
Cited alongside, same era.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D. (2015) · 2015
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R. (2017) · 2017
Cited alongside, same era.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Dann, C., Lattimore, T., and Brunskill, E. (2017) · 2017
Cited alongside, same era.
Few-shot goal inference for visuomotor learning and planning
Xie, A., Singh, A., Levine, S., and Finn, C. (2018) · 2018
Later among the works it cites.
Provably efficient rl with rich observations via latent state decoding
Du, S., Krishnamurthy, A., Jiang, N., Agarwal, A., Dudik, M., and Langford, J. (2019) · 2019
Later among the works it cites.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Zanette, A. and Brunskill, E. (2019) · 2019
Later among the works it cites.
Bandit algorithms
Lattimore, T. and Szepesvári, C. (2020) · 2020
Later among the works it cites.
Navigating to the best policy in markov decision processes
Al Marjani, A., Garivier, A., and Proutiere, A. (2021) · 2021
Closest in time.
Beyond value-function gaps: Improved instance-dependent regret bounds for episodic reinforcement learning
Dann, C., Marinov, T. V., Mohri, M., and Zimmert, J. (2021) · 2021
Closest in time.
Beyond no regret: Instance-dependent pac reinforcement learning
Wagenmaker, A., Simchowitz, M., and Jamieson, K. (2021) · 2021
Closest in time.
Accommodating picky customers: Regret bound and exploration complexity for multi-objective reinforcement learning
Wu, J., Braverman, V., and Yang, L. F. (2021) · 2021
Closest in time.
Fine-grained gap-dependent bounds for tabular mdps via adaptive multi-step bootstrap
Xu, H., Ma, T., and Du, S. S. (2021) · 2021
Closest in time.