Fetching the paper…
Reading the bibliography…
Agents must be able to adapt quickly as an environment changes.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M · 1912
Earlier work this paper cites.
Curiosity in zoo animals
Glickman, S. E. and Sroges, R. W · 1966
Earlier work this paper cites.
A possibility for implementing curiosity and boredom in model-building neural controllers
Schmidhuber, J · 1991
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J · 1992
Earlier work this paper cites.
Prioritized sweeping: Reinforcement learning with less data and less time
Moore, A. W. and Atkeson, C. G · 1993
Earlier work this paper cites.
Lifelong learning algorithms
Thrun, S · 1998
Earlier work this paper cites.
All else being equal be empowered
Klyubin, A. S., Polani, D., and Nehaniv, C. L · 2005
Earlier work this paper cites.
Resilient machines through continuous self-modeling
Bongard, J., Zykov, V., and Lipson, H · 2006
Earlier work this paper cites.
Dealing with non-stationary environments using context detection
Da Silva, B. C., Basso, E. W., Bazzan, A. L., and Engel, P. M · 2006
Earlier work this paper cites.
Bayesian online changepoint detection
Adams, R. P. and MacKay, D. J · 2007
Earlier work this paper cites.
On-line inference for multiple changepoint problems
Fearnhead, P. and Liu, Z · 2007
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
Hunter, J. D · 2007
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Oudeyer, P.-Y., Kaplan, F., and Hafner, V. V · 2007
Earlier work this paper cites.
Data Structures for Statistical Computing in Python
Wes McKinney · 2010
Earlier work this paper cites.
An information-theoretic approach to curiosity-driven reinforcement learning
Still, S. and Precup, D · 2012
Earlier work this paper cites.
Sequential decision-making under non-stationary environments via sequential change-point detection
Hadoux, E., Beynier, A., and Weng, P · 2014
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2015
Earlier work this paper cites.
Pillow (pil fork) documentation, 2015
Clark, A · 2015
Earlier work this paper cites.
Robots that can adapt like animals
Cully, A., Clune, J., Tarapore, D., and Mouret, J.-B · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Variational information maximisation for intrinsically motivated reinforcement learning
Mohamed, S. and Jimenez Rezende, D · 2015
Earlier work this paper cites.
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2015
Earlier work this paper cites.
Incentivizing exploration in reinforcement learning with deep predictive models
Stadie, B. C., Levine, S., and Abbeel, P · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Earlier work this paper cites.
Gregor, K., Rezende, D. J., and Wierstra, D · 2016
Earlier work this paper cites.
Vime: Variational information maximizing exploration
Houthooft, R., Chen, X., Duan, Y., Schulman, J., De Turck, F., and Abbeel, P · 2016
Earlier work this paper cites.
Rusu, A. A., Rabinowitz, N. C., Desjardins, G., Soyer, H., Kirkpatrick, J., Kavukcuoglu, K., Pascanu, R., and Hadsell, R · 2016
Earlier work this paper cites.
Surprise-based intrinsic motivation for deep reinforcement learning
Achiam, J. and Sastry, S · 2017
Earlier work this paper cites.
Continuous adaptation via meta-learning in nonstationary and competitive environments
Al-Shedivat, M., Bansal, T., Burda, Y., Sutskever, I., Mordatch, I., and Abbeel, P · 2017
Earlier work this paper cites.
Quickest change detection approach to optimal control in markov decision processes with model changes
Banerjee, T., Liu, M., and How, J. P · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al · 2017
Earlier work this paper cites.
Learning without forgetting
Li, Z. and Hoiem, D · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Earlier work this paper cites.
# exploration: A study of count-based exploration for deep reinforcement learning
Tang, H., Houthooft, R., Foote, D., Stooke, A., Xi Chen, O., Duan, Y., Schulman, J., DeTurck, F., and Abbeel, P · 2017
Earlier work this paper cites.
Continual learning through synaptic intelligence
Zenke, F., Poole, B., and Ganguli, S · 2017
Earlier work this paper cites.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O · 2018
Earlier work this paper cites.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2018
Cited alongside, same era.
Recurrent world models facilitate policy evolution
Ha, D. and Schmidhuber, J · 2018
Cited alongside, same era.
Learning to play with intrinsically-motivated, self-aware agents
Haber, N., Mrowca, D., Wang, S., Fei-Fei, L. F., and Yamins, D. L · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D · 2018
Cited alongside, same era.
Distributed prioritized experience replay
Horgan, D., Quan, J., Budden, D., Barth-Maron, G., Hessel, M., Van Hasselt, H., and Silver, D · 2018
Cited alongside, same era.
SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python
Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson, A. R. J., Jones, E., Kern, R., Larson, E., Carey, C. J., Polat, İ., Feng, Y., Moore, E. W., VanderPlas, J., Laxalde, D., Perktold, J., Cimrman, R., Henriksen, I., Quintero, E. A., Harris, C. R., Archibald, A. M., Ribeiro, A. H., Pedregosa, F., van Mulbregt, P., and SciPy 1.0 Contributors · 2020
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Yarats, D., Kostrikov, I., and Fergus, R · 2020
Later among the works it cites.
Learning fast adaptation with meta strategy optimization
Yu, W., Tan, J., Bai, Y., Coumans, E., and Ha, S · 2020
Later among the works it cites.
A cell type–specific cortico-subcortical brain circuit for investigatory and novelty-seeking behavior
Ahmadlou, M., Houba, J. H., van Vierbergen, J. F., Giannouli, M., Gimenez, G.-A., van Weeghel, C., Darbanfouladi, M., Shirazi, M. Y., Dziubek, J., Kacem, M., et al · 2021
Later among the works it cites.
Rethinking experience replay: a bag of tricks for continual learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nagabandi, A., Finn, C., and Levine, S · 2018
Cited alongside, same era.
Been there, done that: Meta-learning with episodic recall
Ritter, S., Wang, J., Kurth-Nelson, Z., Jayakumar, S., Blundell, C., Pascanu, R., and Botvinick, M · 2018
Cited alongside, same era.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al · 2018
Cited alongside, same era.
Smirl: Surprise minimizing reinforcement learning in unstable environments
Berseth, G., Geng, D., Devin, C., Rhinehart, N., Finn, C., Jayaraman, D., and Levine, S · 2019
Cited alongside, same era.
Online meta-learning
Finn, C., Rajeswaran, A., Kakade, S., and Levine, S · 2019
Cited alongside, same era.
Meta-learning representations for continual learning
Javed, K. and White, M · 2019
Cited alongside, same era.
Model-based reinforcement learning for atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al · 2019
Cited alongside, same era.
Buzzega, P., Boschini, M., Porrello, A., and Calderara, S · 2021
Later among the works it cites.
Reverb: A framework for experience replay, 2021
Cassirer, A., Barth-Maron, G., Brevdo, E., Ramos, S., Boyd, T., Sottiaux, T., and Kroiss, M · 2021
Later among the works it cites.
Benchmarking the spectrum of agent capabilities
Hafner, D · 2021
Later among the works it cites.
Wilds: A benchmark of in-the-wild distribution shifts
Koh, P. W., Sagawa, S., Marklund, H., Xie, S. M., Zhang, M., Balsubramani, A., Hu, W., Yasunaga, M., Phillips, R. L., Gao, I., et al · 2021
Later among the works it cites.
Rma: Rapid motor adaptation for legged robots
Kumar, A., Fu, Z., Pathak, D., and Malik, J · 2021
Later among the works it cites.
Urlb: Unsupervised reinforcement learning benchmark
Laskin, M., Yarats, D., Liu, H., Lee, K., Zhan, A., Lu, K., Cang, C., Pinto, L., and Abbeel, P · 2021
Later among the works it cites.
Regret minimization experience replay in off-policy reinforcement learning
Liu, X.-H., Xue, Z., Pang, J., Jiang, S., Xu, F., and Yu, Y · 2021
Later among the works it cites.
Discovering and achieving goals via world models
Mendonca, R., Rybkin, O., Daniilidis, K., Hafner, D., and Pathak, D · 2021
Later among the works it cites.
Model-augmented prioritized experience replay
Oh, Y., Shin, J., Yang, E., and Hwang, S. J · 2021
Later among the works it cites.
Interesting object, curious agent: Learning task-agnostic exploration
Parisi, S., Dean, V., Pathak, D., and Gupta, A · 2021
Later among the works it cites.
Extending the wilds benchmark for unsupervised adaptation
Sagawa, S., Koh, P. W., Lee, T., Gao, I., Xie, S. M., Shen, K., Kumar, A., Hu, W., Yasunaga, M., Marklund, H., et al · 2021
Later among the works it cites.
The distracting control suite–a challenging benchmark for reinforcement learning from pixels
Stone, A., Ramirez, O., Konolige, K., and Jonschkowski, R · 2021
Later among the works it cites.
seaborn: statistical data visualization
Waskom, M. L · 2021
Later among the works it cites.
Deep reinforcement learning amidst continual structured non-stationarity
Xie, A., Harrison, J., and Finn, C · 2021
Later among the works it cites.
Reinforcement learning with prototypical representations
Yarats, D., Fergus, R., Lazaric, A., and Pinto, L · 2021
Later among the works it cites.
Learning to prioritize planning updates in model-based reinforcement learning
Burega, B., Martin, J. D., and Bowling, M · 2022
Later among the works it cites.
Transdreamer: Reinforcement learning with transformer world models
Chen, C., Wu, Y.-F., Yoon, J., and Ahn, S · 2022
Later among the works it cites.
Dreamerpro: Reconstruction-free model-based reinforcement learning with prototypical representations
Deng, F., Jang, I., and Ahn, S · 2022
Later among the works it cites.
Sample-efficient reinforcement learning by breaking the replay ratio barrier
D’Oro, P., Schwarzer, M., Nikishin, E., Bacon, P.-L., Bellemare, M. G., and Courville, A · 2022
Later among the works it cites.
Byol-explore: Exploration by bootstrapped prediction
Guo, Z. D., Thakoor, S., Pîslar, M., Pires, B. A., Altché, F., Tallec, C., Saade, A., Calandriello, D., Grill, J.-B., Tang, Y., et al · 2022
Later among the works it cites.
Deep hierarchical planning from pixels
Hafner, D., Lee, K.-H., Fischer, I., and Abbeel, P · 2022
Later among the works it cites.
Towards continual reinforcement learning: A review and perspectives
Khetarpal, K., Riemer, M., Rish, I., and Precup, D · 2022
Later among the works it cites.
Transformers are sample efficient world models
Micheli, V., Alonso, E., and Fleuret, F · 2022
Later among the works it cites.
Understanding and mitigating the limitations of prioritized replay
Pan, Y., Mei, J., Farahmand, A.-m., White, M., Yao, H., Rohani, M., and Luo, J · 2022
Later among the works it cites.
Experience replay with likelihood-free importance weights
Sinha, S., Song, J., Garg, A., and Ermon, S · 2022
Later among the works it cites.
Towards evaluating adaptivity of model-based reinforcement learning methods
Wan, Y., Rahimi-Kalahroudi, A., Rajendran, J., Momennejad, I., Chandar, S., and Van Seijen, H. H · 2022
Later among the works it cites.
Daydreamer: World models for physical robot learning
Wu, P., Escontrela, A., Hafner, D., Goldberg, K., and Abbeel, P · 2022
Later among the works it cites.
Neuro-symbolic world models for adapting to open world novelty
Balloch, J., Lin, Z., Wright, R., Peng, X., Hussain, M., Srinivas, A., Kim, J., and Riedl, M. O · 2023
Closest in time.
The role of experience in prioritizing hippocampal replay
Gorriz, M. H., Takigawa, M., and Bendor, D · 2023
Closest in time.
Mastering diverse domains through world models
Hafner, D., Pasukonis, J., Ba, J., and Lillicrap, T · 2023
Closest in time.
Model-based reinforcement learning: A survey
Moerland, T. M., Broekens, J., Plaat, A., Jonker, C. M., et al · 2023
Closest in time.
Human-timescale adaptation in an open-ended task space
Team, A. A., Bauer, J., Baumli, K., Baveja, S., Behbahani, F., Bhoopchand, A., Bradley-Schmieg, N., Chang, M., Clay, N., Collister, A., et al · 2023
Closest in time.