Fetching the paper…
Reading the bibliography…
Efficient exploration is one of the main challenges in reinforcement learning (RL).
Go-explore: a new approach for hard-exploration problems
Ecoffet, A., Huizinga, J., Lehman, J., Stanley, K. O., and Clune, J. (2019) · 1901
Earlier work this paper cites.
Zanette, A. and Brunskill, E. (2019) · 1901
Earlier work this paper cites.
Improvisation through physical understanding: Using novel objects as tools with visual foresight
Xie, A., Ebert, F., Levine, S., and Finn, C. (2019) · 1904
Earlier work this paper cites.
Exact robot navigation using artificial potential functions
Rimon, E. and Koditschek, D. E. (1992) · 1992
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Dietterich, T. G. (2000) · 2000
Earlier work this paper cites.
Reward-free exploration for reinforcement learning
Jin, C., Krishnamurthy, A., Simchowitz, M., and Yu, T. (2020) · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S. (2002) · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Kakade, S. M. et al. (2003) · 2003
Earlier work this paper cites.
The sample complexity of exploration in the multi-armed bandit problem
Mannor, S. and Tsitsiklis, J. N. (2004) · 2004
Earlier work this paper cites.
Almost optimal model-free reinforcement learning via reference-advantage decomposition
Zhang, Z., Zhou, Y., and Ji, X. (2020) · 2004
Earlier work this paper cites.
Pac model-free reinforcement learning
Strehl, A. L., Li, L., Wiewiora, E., Langford, J., and Littman, M. L. (2006) · 2006
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P. (2010) · 2010
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
Dann, C. and Brunskill, E. (2015) · 2015
Cited alongside, same era.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D. (2015) · 2015
Cited alongside, same era.
Socially compliant mobile robot navigation via inverse reinforcement learning
Kretzschmar, H., Spies, M., Sprunk, C., and Burgard, W. (2016) · 2016
Cited alongside, same era.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Agrawal, S. and Jia, R. (2017) · 2017
Cited alongside, same era.
Zero-shot task generalization with multi-task deep reinforcement learning
Oh, J., Singh, S., Lee, H., and Kohli, P. (2017) · 2017
Later among the works it cites.
A deep hierarchical approach to lifelong learning in minecraft
Tessler, C., Givony, S., Zahavy, T., Mankowitz, D. J., and Mannor, S. (2017) · 2017
Later among the works it cites.
Policy certificates: Towards accountable reinforcement learning
Dann, C., Li, L., Wei, W., and Brunskill, E. (2018) · 2018
Later among the works it cites.
Notes on tabular methods
Jiang, N. (2018) · 2018
Later among the works it cites.
Is q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I. (2018) · 2018
Later among the works it cites.
Learning by playing-solving sparse reward tasks from scratch
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, O. P., and Zaremba, W. (2017) · 2017
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R. (2017) · 2017
Cited alongside, same era.
Deep visual foresight for planning robot motion
Finn, C. and Levine, S. (2017) · 2017
Cited alongside, same era.
Riedmiller, M., Hafner, R., Lampe, T., Neunert, M., Degrave, J., Van de Wiele, T., Mnih, V., Heess, N., and Springenberg, J. T. (2018) · 2018
Later among the works it cites.
Meta reinforcement learning with latent variable gaussian processes
Sæmundsson, S., Hofmann, K., and Deisenroth, M. P. (2018) · 2018
Later among the works it cites.
Few-shot goal inference for visuomotor learning and planning
Xie, A., Singh, A., Levine, S., and Finn, C. (2018) · 2018
Later among the works it cites.
Drn: A deep reinforcement learning framework for news recommendation
Zheng, G., Zhang, F., Zheng, Z., Xiang, Y., Yuan, N. J., Xie, X., and Li, Z. (2018) · 2018
Later among the works it cites.