Fetching the paper…
Reading the bibliography…
We introduce a new unsupervised pretraining objective for reinforcement learning.
Improving generalization for temporal difference learning: The successor representation
Dayan, P · 1993
Earlier work this paper cites.
Nonparametric entropy estimation: An overview
Beirlant, J · 1997
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
The im algorithm: A variational approach to information maximization
Barber, D. and Agakov, F. V · 2003
Earlier work this paper cites.
Nearest neighbor estimates of entropy
Singh, H., Misra, N., Hnizdo, V., Fedorowicz, A., and Demchuk, E · 2003
Earlier work this paper cites.
Intrinsically motivated learning of hierarchical collections of skills
Barto, A. G., Singh, S., and Chentanez, N · 2004
Earlier work this paper cites.
An intrinsic reward mechanism for efficient exploration
Simsek, Ö. and Barto, A. G · 2006
Earlier work this paper cites.
What is intrinsic motivation? a typology of computational approaches
Oudeyer, P.-Y. and Kaplan, F · 2009
Earlier work this paper cites.
Intrinsic motivation and reinforcement learning
Barto, A. G · 2013
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Variational information maximisation for intrinsically motivated reinforcement learning
Mohamed, S. and Rezende, D. J · 2015
Earlier work this paper cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Clevert, D., Unterthiner, T., and Hochreiter, S · 2016
Earlier work this paper cites.
Gregor, K., Rezende, D. J., and Wierstra, D · 2016
Earlier work this paper cites.
VIME: variational information maximizing exploration
Houthooft, R., Chen, X., Duan, Y., Schulman, J., Turck, F. D., and Abbeel, P · 2016
Earlier work this paper cites.
Deep successor reinforcement learning
Kulkarni, T. D., Saeedi, A., Gautam, S., and Gershman, S. J · 2016
Earlier work this paper cites.
Surprise-based intrinsic motivation for deep reinforcement learning
Achiam, J. and Sastry, S · 2017
Earlier work this paper cites.
Hindsight experience replay
Andrychowicz, M., Crow, D., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, P., and Zaremba, W · 2017
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., Silver, D., and van Hasselt, H · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Cited alongside, same era.
The pycolab game engine
Stepleton, T · 2017
Cited alongside, same era.
Transfer in deep reinforcement learning using successor features and generalised policy improvement
Barreto, A., Borsa, D., Quan, J., Schaul, T., Silver, D., Hessel, M., Mankowitz, D. J., Zídek, A., and Munos, R · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M. G., and Silver, D · 2018
Cited alongside, same era.
Deep contextualized word representations
Peters, M., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L · 2018
Fast reinforcement learning with generalized policy updates
Barreto, A., Hou, S., Borsa, D., Silver, D., and Precup, D · 2020
Later among the works it cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Later among the works it cites.
Explore, discover and learn: Unsupervised discovery of state-covering skills
Campos, V., Trott, A., Xiong, C., Socher, R., Giró-i-Nieto, X., and Torres, J · 2020
Later among the works it cites.
A simple framework for contrastive learning of visual representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. E · 2020
Later among the works it cites.
Big self-supervised models are strong semi-supervised learners
Chen, T., Kornblith, S., Swersky, K., Norouzi, M., and Hinton, G. E · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Solving rubik’s cube with a robot hand
Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., et al · 2019
Cited alongside, same era.
Universal successor features approximators
Borsa, D., Barreto, A., Quan, J., Mankowitz, D. J., van Hasselt, H., Munos, R., Silver, D., and Schaul, T · 2019
Cited alongside, same era.
Large-scale study of curiosity-driven learning
Burda, Y., Edwards, H., Pathak, D., Storkey, A. J., Darrell, T., and Efros, A. A · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2019
Cited alongside, same era.
Provably efficient maximum entropy exploration
Hazan, E., Kakade, S. M., Singh, K., and Soest, A. V · 2019
Cited alongside, same era.
Fast task inference with variational intrinsic successor features
Hansen, S., Dabney, W., Barreto, A., Warde-Farley, D., de Wiele, T. V., and Mnih, V · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. B · 2020
Later among the works it cites.
Model based reinforcement learning for atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., Mohiuddin, A., Sepassi, R., Tucker, G., and Michalewski, H · 2020
Later among the works it cites.
Do recent advancements in model-based deep reinforcement learning really improve data efficiency?
Kielak, K · 2020
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Kostrikov, I., Yarats, D., and Fergus, R · 2020
Later among the works it cites.
CURL: contrastive unsupervised representations for reinforcement learning
Laskin, M., Srinivas, A., and Abbeel, P · 2020
Later among the works it cites.
Count-based exploration with the successor representation
Machado, M. C., Bellemare, M. G., and Bowling, M · 2020
Later among the works it cites.
A policy gradient method for task-agnostic exploration
Mutti, M., Pratissoli, L., and Restelli, M · 2020
Later among the works it cites.
Dynamics-aware unsupervised discovery of skills
Sharma, A., Gu, S., Levine, S., Kumar, V., and Hausman, K · 2020
Later among the works it cites.
Coverage as a principle for discovering transferable behavior in reinforcement learning, 2021
Campos, V., Sprechmann, P., Hansen, S. S., Barreto, A., Blundell, C., Vitvitskyi, A., Kapturowski, S., and Badia, A. P · 2021
Closest in time.
Behavior from the void: Unsupervised active pre-training
Liu, H. and Abbeel, P · 2021
Closest in time.
Data-efficient reinforcement learning with self-predictive representations
Schwarzer, M., Anand, A., Goel, R., Hjelm, R. D., Courville, A., and Bachman, P · 2021
Closest in time.
State entropy maximization with random encoders for efficient exploration
Seo, Y., Chen, L., Shin, J., Lee, H., Abbeel, P., and Lee, K · 2021
Closest in time.
Reinforcement learning with prototypical representations
Yarats, D., Fergus, R., Lazaric, A., and Pinto, L · 2021
Closest in time.