Fetching the paper…
Reading the bibliography…
Most deep reinforcement learning (RL) algorithms distill experience into parametric behavior policies or value functions via gradient updates.
Infobot: Transfer and exploration via the information bottleneck
Goyal, A., Islam, R., Strouse, D., Ahmed, Z., Botvinick, M., Larochelle, H., Bengio, Y., and Levine, S · 1901
Earlier work this paper cites.
Reinforcement learning with competitive ensembles of information-constrained primitives
Goyal, A., Sodhani, S., Binas, J., Peng, X. B., Levine, S., and Bengio, Y · 1906
Earlier work this paper cites.
Recurrent independent mechanisms
Goyal, A., Lamb, A., Hoffmann, J., Sodhani, S., Levine, S., Bengio, Y., and Schölkopf, B · 1909
Earlier work this paper cites.
Planning using a temporal world model
Allen, J. F. and Koomen, J. A · 1983
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Sutton, R. S · 1991
Earlier work this paper cites.
An introduction to case-based reasoning
Kolodner, J. L · 1992
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, L.-J · 1992
Earlier work this paper cites.
Case-based reasoning: experiences, lessons, and future directions
Leake, D. B · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Elements of information theory
Cover, T. M · 1999
Earlier work this paper cites.
The information bottleneck method
Tishby, N., Pereira, F. C., and Bialek, W · 2000
Earlier work this paper cites.
The variational bandwidth bottleneck: Stochastic evaluation on an information budget
Goyal, A., Bengio, Y., Botvinick, M., and Levine, S · 2004
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., Geurts, P., and Wehenkel, L · 2005
Earlier work this paper cites.
Neural fitted Q iteration–first experiences with a data efficient neural reinforcement learning method
Riedmiller, M · 2005
Earlier work this paper cites.
Big self-supervised models are strong semi-supervised learners
Chen, T., Kornblith, S., Swersky, K., Norouzi, M., and Hinton, G · 2006
Earlier work this paper cites.
Object files and schemata: Factorizing declarative and procedural knowledge in dynamical systems
Goyal, A., Lamb, A., Gampa, P., Beaudoin, P., Levine, S., Blundell, C., Bengio, Y., and Mozer, M · 2006
Earlier work this paper cites.
Rl unplugged: A suite of benchmarks for offline reinforcement learning
Gulcehre, C., Wang, Z., Novikov, A., Paine, T. L., Colmenarejo, S. G., Zolna, K., Agarwal, R., Merel, J., Mankowitz, D., Paduraru, C., et al · 2006
Earlier work this paper cites.
Sample-based learning and search with permanent and transient memories
Silver, D., Sutton, R. S., and Müller, M · 2008
Earlier work this paper cites.
Reinforcement learning and simulation-based search in computer go
Silver, D · 2009
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Auto-encoding variational Bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, D. J., Mohamed, S., and Wierstra, D · 2014
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients, 2015
Heess, N., Wayne, G., Silver, D., Lillicrap, T., Tassa, Y., and Erez, T · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Deep variational information bottleneck
Alemi, A. A., Fischer, I., Dillon, J. V., and Murphy, K · 2016
Earlier work this paper cites.
Interaction networks for learning about objects, relations and physics
Battaglia, P. W., Pascanu, R., Lai, M., Rezende, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Blundell, C., Uria, B., Pritzel, A., Li, Y., Ruderman, A., Leibo, J. Z., Rae, J., Wierstra, D., and Hassabis, D · 2016
Earlier work this paper cites.
Learning and transfer of modulated locomotor controllers
Heess, N., Wayne, G., Tassa, Y., Lillicrap, T., Riedmiller, M., and Silver, D · 2016
Earlier work this paper cites.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Cited alongside, same era.
Stochastic neural networks for hierarchical reinforcement learning
Florensa, C., Duan, Y., and Abbeel, P · 2017
Cited alongside, same era.
Meta learning shared hierarchies
Frans, K., Ho, J., Chen, X., Abbeel, P., and Schulman, J · 2017
Cited alongside, same era.
Building machines that learn and think like people
Lake, B. M., Ullman, T. D., Tenenbaum, J. B., and Gershman, S. J · 2017
Cited alongside, same era.
Learning model-based planning from scratch
Pascanu, R., Li, Y., Vinyals, O., Heess, N., Buesing, L., Racanière, S., Reichert, D., Weber, T., Wierstra, D., and Battaglia, P · 2017
Cited alongside, same era.
Generalization of reinforcement learners with working and episodic memory
Fortunato, M., Tan, M., Faulkner, R., Hansen, S., Badia, A. P., Buttimore, G., Deck, C., Leibo, J. Z., and Blundell, C · 2019
Later among the works it cites.
Information asymmetry in KL-regularized RL
Galashov, A., Jayakumar, S. M., Hasenclever, L., Tirumala, D., Schwarz, J., Desjardins, G., Czarnecki, W. M., Teh, Y. W., Pascanu, R., and Heess, N · 2019
Later among the works it cites.
Learning dynamics model in reinforcement learning by incorporating the long term future
Ke, N. R., Singh, A., Touati, A., Goyal, A., Bengio, Y., Parikh, D., and Batra, D · 2019
Later among the works it cites.
Latent retrieval for weakly supervised open domain question answering
Lee, K., Chang, M.-W., and Toutanova, K · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neural episodic control
Pritzel, A., Uria, B., Srinivasan, S., Badia, A. P., Vinyals, O., Hassabis, D., Wierstra, D., and Blundell, C · 2017
Cited alongside, same era.
Imagination-augmented agents for deep reinforcement learning
Racanière, S., Weber, T., Reichert, D. P., Buesing, L., Guez, A., Rezende, D., Badia, A. P., Vinyals, O., Heess, N., Li, Y., et al · 2017
Cited alongside, same era.
The predictron: End-to-end learning and planning
Silver, D., Hasselt, H., Hessel, M., Schaul, T., Guez, A., Harley, T., Dulac-Arnold, G., Reichert, D., Rabinowitz, N., Barreto, A., et al · 2017
Cited alongside, same era.
Distral: Robust multitask reinforcement learning
Teh, Y., Bapst, V., Czarnecki, W. M., Quan, J., Kirkpatrick, J., Hadsell, R., Heess, N., and Pascanu, R · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
Vecerik, M., Hester, T., Scholz, J., Wang, F., Pietquin, O., Piot, B., Heess, N., Rothörl, T., Lampe, T., and Riedmiller, M · 2017
Cited alongside, same era.
Feudal networks for hierarchical reinforcement learning
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K · 2017
Cited alongside, same era.
Regularization matters in policy optimization
Liu, Z., Li, X., Kang, B., and Darrell, T · 2019
Later among the works it cites.
Hierarchical motor control in mammals and machines
Merel, J., Botvinick, M., and Wayne, G · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al · 2019
Later among the works it cites.
When to use parametric models in reinforcement learning?
van Hasselt, H. P., Hessel, M., and Aslanides, J · 2019
Later among the works it cites.
Grandmaster level in starcraft II using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Later among the works it cites.
Watters, N., Matthey, L., Bosnjak, M., Burgess, C. P., and Lerchner, A · 2019
Later among the works it cites.
Causalworld: A robotic manipulation benchmark for causal structure and transfer learning
Ahmed, O., Träuble, F., Goyal, A., Neitz, A., Bengio, Y., Schölkopf, B., Wüthrich, M., and Bauer, S · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Later among the works it cites.
REALM: Retrieval-augmented language model pre-training
Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M.-W · 2020
Later among the works it cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al · 2020
Later among the works it cites.
Deep reinforcement and infomax learning
Mazoure, B., Combes, R. T. d., Doan, T., Bachman, P., and Hjelm, R. D · 2020
Later among the works it cites.
AWAC: Accelerating online reinforcement learning with offline datasets
Nair, A., Dalal, M., Gupta, A., and Levine, S · 2020
Later among the works it cites.
Accelerating reinforcement learning with learned skill priors
Pertsch, K., Lee, Y., and Lim, J. J · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al · 2020
Later among the works it cites.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
Siegel, N. Y., Springenberg, J. T., Berkenkamp, F., Abdolmaleki, A., Neunert, M., Lampe, T., Hafner, R., Heess, N., and Riedmiller, M · 2020
Later among the works it cites.
Local search for policy iteration in continuous control
Springenberg, J. T., Heess, N., Mankowitz, D., Merel, J., Byravan, A., Abdolmaleki, A., Kay, J., Degrave, J., Schrittwieser, J., Tassa, Y., et al · 2020
Later among the works it cites.
Entity abstraction in visual model-based reinforcement learning
Veerapaneni, R., Co-Reyes, J. D., Chang, M., Janner, M., Finn, C., Wu, J., Tenenbaum, J., and Levine, S · 2020
Later among the works it cites.
Episodic reinforcement learning with associative memory
Zhu, G., Lin, Z., Yang, G., and Zhang, C · 2020
Later among the works it cites.
Coberl: Contrastive bert for reinforcement learning
Banino, A., Badia, A. P., Walker, J. C., Scholtes, T., Mitrovic, J., and Blundell, C · 2021
Later among the works it cites.
Improving language models by retrieving from trillions of tokens
Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., Driessche, G. v. d., Lespiau, J.-B., Damoc, B., Clark, A., et al · 2021
Later among the works it cites.
From motor control to team play in simulated humanoid football
Liu, S., Lever, G., Wang, Z., Merel, J., Eslami, S., Hennes, D., Czarnecki, W. M., Tassa, Y., Omidshafiei, S., Abdolmaleki, A., et al · 2021
Later among the works it cites.
The episodic flanker effect: Memory retrieval as attention turned inward
Logan, G. D., Cox, G. E., Annis, J., and Lindsey, D. R · 2021
Later among the works it cites.
Ernie 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation
Sun, Y., Wang, S., Feng, S., Ding, S., Pang, C., Shang, J., Liu, J., Chen, X., Zhao, Y., Lu, Y., et al · 2021
Later among the works it cites.
Exponential lower bounds for planning in mdps with linearly-realizable optimal action-value functions
Weisz, G., Amortila, P., and Szepesvári, C · 2021
Later among the works it cites.