Fetching the paper…
Reading the bibliography…
In the maximum state entropy exploration framework, an agent interacts with a reward-free environment to learn a policy that maximizes the entropy of the expected state visitations it is inducing.
Cheung, W. C · 1905
Earlier work this paper cites.
Optimal control of Markov decision processes with incomplete state estimation
Astrom, K. J · 1965
Earlier work this paper cites.
The complexity of Markov decision processes
Papadimitriou, C. H. and Tsitsiklis, J. N · 1987
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
Williams, R. J. and Zipser, D · 1989
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P · 1992
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P., Littman, M. L., and Cassandra, A. R · 1998
Earlier work this paper cites.
Complexity of finite-horizon Markov decision process problems
Mundhenk, M., Goldsmith, J., Lusena, C., and Allender, E · 2000
Earlier work this paper cites.
Nonapproximability results for partially observable Markov decision processes
Lusena, C., Goldsmith, J., and Mundhenk, M · 2001
Earlier work this paper cites.
Introduction to probability
Bertsekas, D. P. and Tsitsiklis, J. N · 2002
Earlier work this paper cites.
Bandit based Monte-Carlo planning
Kocsis, L. and Szepesvári, C · 2006
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
Peters, J. and Schaal, S · 2008
Earlier work this paper cites.
An analysis of model-based interval estimation for Markov decision processes
Strehl, A. L. and Littman, M. L · 2008
Earlier work this paper cites.
Computational complexity: a modern approach
Arora, S. and Barak, B · 2009
Earlier work this paper cites.
The difficulty of training deep architectures and the effect of unsupervised pre-training
Erhan, D., Manzagol, P.-A., Bengio, Y., Bengio, S., and Vincent, P · 2009
Cited alongside, same era.
Why does unsupervised pre-training help deep learning?
Erhan, D., Courville, A., Bengio, Y., and Vincent, P · 2010
Cited alongside, same era.
The steady-state control problem for Markov decision processes
Akshay, S., Bertrand, N., Haddad, S., and Helouet, L · 2013
Cited alongside, same era.
A survey on policy search for robotics
Deisenroth, M. P., Neumann, G., Peters, J., et al · 2013
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Cited alongside, same era.
Variational policy gradient method for reinforcement learning with general utilities
Zhang, J., Koppel, A., Bedi, A. S., Szepesvari, C., and Wang, M · 2020
Later among the works it cites.
Coverage as a principle for discovering transferable behavior in reinforcement learning
Campos, V., Sprechmann, P., Hansen, S., Barreto, A., Kapturowski, S., Vitvitskyi, A., Badia, A. P., and Blundell, C · 2021
Later among the works it cites.
Geometric entropic exploration
Guo, Z. D., Azar, M. G., Saade, A., Thakoor, S., Piot, B., Pires, B. A., Valko, M., Mesnard, T., Lattimore, T., and Munos, R · 2021
Later among the works it cites.
RL for latent MDPs: Regret guarantees and a lower bound
Kwon, J., Efroni, Y., Caramanis, C., and Mannor, S · 2021
Later among the works it cites.
URLB: Unsupervised reinforcement learning benchmark
Laskin, M., Yarats, D., Liu, H., Lee, K., Zhan, A., Lu, K., Cang, C., Pinto, L., and Abbeel, P · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hallak, A., Di Castro, D., and Mannor, S · 2015
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Provably efficient maximum entropy exploration
Hazan, E., Kakade, S., Singh, K., and Van Soest, A · 2019
Cited alongside, same era.
Efficient exploration via state marginal matching
Lee, L., Eysenbach, B., Parisotto, E., Xing, E., Levine, S., and Salakhutdinov, R · 2019
Cited alongside, same era.
Active exploration in Markov decision processes
Tarbouriech, J. and Lazaric, A · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Reward-free exploration for reinforcement learning
Jin, C., Krishnamurthy, A., Simchowitz, M., and Yu, T · 2020
Cited alongside, same era.
Later among the works it cites.
Task-agnostic exploration via policy gradient of a non-parametric state entropy estimate
Mutti, M., Pratissoli, L., and Restelli, M · 2021
Later among the works it cites.
State entropy maximization with random encoders for efficient exploration
Seo, Y., Chen, L., Shin, J., Lee, H., Abbeel, P., and Lee, K · 2021
Later among the works it cites.
A provably efficient sample collection strategy for reinforcement learning
Tarbouriech, J., Pirotta, M., Valko, M., and Lazaric, A · 2021
Later among the works it cites.
Reinforcement learning with prototypical representations
Yarats, D., Fergus, R., Lazaric, A., and Pinto, L · 2021
Later among the works it cites.
Exploration by maximizing Rényi entropy for reward-free RL framework
Zhang, C., Cai, Y., Huang, L., and Li, J · 2021
Later among the works it cites.
Non-Markovian policies occupancy measures
Laroche, R., Combes, R. T. d., and Buckman, J · 2022
Closest in time.
Unsupervised reinforcement learning in multiple environments
Mutti, M., Mancassola, M., and Restelli, M · 2022
Closest in time.
k-means maximum entropy exploration
Nedergaard, A. and Cook, M · 2022
Closest in time.