Fetching the paper…
Reading the bibliography…
Reinforcement learning algorithms struggle when the reward signal is very sparse.
Asymptotic evaluation of certain markov process expectations for large time. iv
Donsker, M. D. and Varadhan, S. S · 1983
Earlier work this paper cites.
Representation of information in spatial maps which are produced by self-organization
Kohonen, T · 1983
Earlier work this paper cites.
Self-organizing maps of symbol strings
Kohonen, T. and Somervuo, P · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S · 1999
Earlier work this paper cites.
Robust principal component analysis?
Candès, E. J., Li, X., Ma, Y., and Wright, J · 2011
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Skip context tree switching
Bellemare, M., Veness, J., and Talvitie, E · 2014
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. L · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Variational information maximisation for intrinsically motivated reinforcement learning
Mohamed, S. and Rezende, D. J · 2015
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
Oh, J., Guo, X., Lee, H., Lewis, R. L., and Singh, S · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
Stadie, B. C., Levine, S., and Abbeel, P · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Cited alongside, same era.
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
Chen, X., Duan, Y., Houthooft, R., Schulman, J., Sutskever, I., and Abbeel, P · 2016
Cited alongside, same era.
Ex2: Exploration with exemplar models for deep reinforcement learning
Fu, J., Co-Reyes, J., and Levine, S · 2017
Later among the works it cites.
Count-based exploration with neural density models
Ostrovski, G., Bellemare, M. G., Oord, A. v. d., and Munos, R · 2017
Later among the works it cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
# exploration: A study of count-based exploration for deep reinforcement learning
Tang, H., Houthooft, R., Foote, D., Stooke, A., Chen, X., Duan, Y., Schulman, J., DeTurck, F., and Abbeel, P · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Duan, Y., Chen, X., Houthooft, R., Schulman, J., and Abbeel, P · 2016
Cited alongside, same era.
Vime: Variational information maximizing exploration
Houthooft, R., Chen, X., Duan, Y., Schulman, J., De Turck, F., and Abbeel, P · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
f-gan: Training generative neural samplers using variational divergence minimization
Nowozin, S., Cseke, B., and Tomioka, R · 2016
Cited alongside, same era.
Conditional image generation with pixelcnn decoders
van den Oord, A., Kalchbrenner, N., Espeholt, L., Vinyals, O., Graves, A., et al · 2016
Cited alongside, same era.
Independently controllable features
Bengio, E., Thomas, V., Pineau, J., Precup, D., and Bengio, Y · 2017
Cited alongside, same era.
Disentangling the independently controllable factors of variation by interacting with the world
Thomas, V., Bengio, E., Fedus, W., Pondard, J., Beaudoin, P., Larochelle, H., Pineau, J., Precup, D., and Bengio, Y
Cited in the paper.
Mutual information neural estimation
Belghazi, I., Rajeswar, S., Baratin, A., Hjelm, R. D., and Courville, A · 2018
Closest in time.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O · 2018
Closest in time.
Gep-pg: Decoupling exploration and exploitation in deep reinforcement learning algorithms
Colas, C., Sigaud, O., and Oudeyer, P.-Y · 2018
Closest in time.
Learning deep representations by mutual information estimation and maximization
Hjelm, R. D., Fedorov, A., Lavoie-Marchildon, S., Grewal, K., Trischler, A., and Bengio, Y · 2018
Closest in time.
Near-optimal representation learning for hierarchical reinforcement learning
Nachum, O., Gu, S., Lee, H., and Levine, S · 2018
Closest in time.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Closest in time.