Safe model-based reinforcement learning with stability guarantees
Berkenkamp, F., Turchetta, M., Schoellig, A., and Krause, A. (2017) · 2017
Later among the works it cites.
Recurrent environment simulators
Original
Chiappa, S., Racaniere, S., Wierstra, D., and Mohamed, S. (2017) · 2017
Later among the works it cites.
Stochastic neural networks for hierarchical reinforcement learning
Original
Florensa, C., Duan, Y., and Abbeel, P. (2017) · 2017
Later among the works it cites.
Multi-level discovery of deep options
Original
Fox, R., Krishnan, S., Stoica, I., and Goldberg, K. (2017) · 2017
Later among the works it cites.
A disentangled recognition and nonlinear dynamics model for unsupervised learning
Fraccaro, M., Kamronn, S., Paquet, U., and Winther, O. (2017) · 2017
Later among the works it cites.
Generative temporal models with memory
Original
Gemici, M., Hung, C.-C., Santoro, A., Wayne, G., Mohamed, S., Rezende, D. J., Amos, D., and Lillicrap, T. (2017) · 2017
Later among the works it cites.
Reinforcement learning and episodic memory in humans and animals: an integrative framework
Gershman, S. J. and Daw, N. D. (2017) · 2017
Later among the works it cites.
Metacontrol for adaptive imagination-based optimization
Original
Hamrick, J. B., Ballard, A. J., Pascanu, R., Vinyals, O., Heess, N., and Battaglia, P. W. (2017) · 2017
Later among the works it cites.
Hierarchical reinforcement learning
Hengst, B. (2017) · 2017
Later among the works it cites.
Video pixel networks
Kalchbrenner, N., van den Oord, A., Simonyan, K., Danihelka, I., Vinyals, O., Graves, A., and Kavukcuoglu, K. (2017) · 2017
Later among the works it cites.
Uncertainty-driven imagination for continuous deep reinforcement learning
Kalweit, G. and Boedecker, J. (2017) · 2017
Later among the works it cites.
Data-efficient reinforcement learning with probabilistic model predictive control
Original
Kamthe, S. and Deisenroth, M. P. (2017) · 2017
Later among the works it cites.
Schema networks: Zero-shot transfer with a generative causal model of intuitive physics
Kansky, K., Silver, T., Mély, D. A., Eldawy, M., Lázaro-Gredilla, M., Lou, X., Dorfman, N., Sidor, S., Phoenix, S., and George, D. (2017) · 2017
Later among the works it cites.
A laplacian framework for option discovery in reinforcement learning
Machado, M. C., Bellemare, M. G., and Bowling, M. (2017) · 2017
Later among the works it cites.
Teacher-student curriculum learning
Original
Matiisen, T., Oliver, A., Cohen, T., and Schulman, J. (2017) · 2017
Later among the works it cites.
Prediction and control with temporal segment models
Mishra, N., Abbeel, P., and Mordatch, I. (2017) · 2017
Later among the works it cites.
The successor representation in human reinforcement learning
Momennejad, I., Russek, E. M., Cheong, J. H., Botvinick, M. M., Daw, N. D., and Gershman, S. J. (2017) · 2017
Later among the works it cites.
Value prediction network
Oh, J., Singh, S., and Lee, H. (2017) · 2017
Later among the works it cites.
Why is posterior sampling better than optimism for reinforcement learning?
Osband, I. and Van Roy, B. (2017) · 2017
Later among the works it cites.
Count-based exploration with neural density models
Ostrovski, G., Bellemare, M. G., van den Oord, A., and Munos, R. (2017) · 2017
Later among the works it cites.
Learning model-based planning from scratch
Original
Pascanu, R., Li, Y., Vinyals, O., Heess, N., Buesing, L., Racanière, S., Reichert, D., Weber, T., Wierstra, D., and Battaglia, P. (2017) · 2017
Later among the works it cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T. (2017) · 2017
Later among the works it cites.
Parameter space noise for exploration
Original
Plappert, M., Houthooft, R., Dhariwal, P., Sidor, S., Chen, R. Y., Chen, X., Asfour, T., Abbeel, P., and Andrychowicz, M. (2017) · 2017
Later among the works it cites.
Survey of model-based reinforcement learning: Applications on robotics
Polydoros, A. S. and Nalpantidis, L. (2017) · 2017
Later among the works it cites.
Neural Episodic Control
Pritzel, A., Uria, B., Srinivasan, S., Badia, A. P., Vinyals, O., Hassabis, D., Wierstra, D., and Blundell, C. (2017) · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Racanière, S., Weber, T., Reichert, D., Buesing, L., Guez, A., Rezende, D. J., Badia, A. P., Vinyals, O., Heess, N., Li, Y., et al. (2017) · 2017
Later among the works it cites.
Multi-objective decision making
Roijers, D. M. and Whiteson, S. (2017) · 2017
Later among the works it cites.
Hierarchical and interpretable skill acquisition in multi-task reinforcement learning
Original
Shu, T., Xiong, C., and Socher, R. (2017) · 2017
Later among the works it cites.
Self-correcting models for model-based reinforcement learning
Talvitie, E. (2017) · 2017
Later among the works it cites.
A deep hierarchical approach to lifelong learning in minecraft
Tessler, C., Givony, S., Zahavy, T., Mankowitz, D. J., and Mannor, S. (2017) · 2017
Later among the works it cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., and Abbeel, P. (2017) · 2017
Later among the works it cites.
Feudal networks for hierarchical reinforcement learning
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K. (2017) · 2017
Later among the works it cites.
Seqgan: Sequence generative adversarial nets with policy gradient
Yu, L., Zhang, W., Wang, J., and Yu, Y. (2017) · 2017
Later among the works it cites.
Variational option discovery algorithms
Original
Achiam, J., Edwards, H., Amodei, D., and Abbeel, P. (2018) · 2018
Later among the works it cites.
Towards a Simple Approach to Multi-step Model-based Reinforcement Learning
Original
Asadi, K., Cater, E., Misra, D., and Littman, M. L. (2018) · 2018
Later among the works it cites.
Relational inductive biases, deep learning, and graph networks
Original
Battaglia, P. W., Hamrick, J. B., Bapst, V., Sanchez-Gonzalez, A., Zambaldi, V., Malinowski, M., Tacchetti, A., Raposo, D., Santoro, A., Faulkner, R., et al. (2018) · 2018
Later among the works it cites.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Buckman, J., Hafner, D., Tucker, G., Brevdo, E., and Lee, H. (2018) · 2018
Later among the works it cites.
Learning and querying fast generative models for reinforcement learning
Original
Buesing, L., Weber, T., Racaniere, S., Eslami, S., Rezende, D., Reichert, D. P., Viola, F., Besse, F., Gregor, K., Hassabis, D., et al. (2018) · 2018
Later among the works it cites.
Contingency-aware exploration in reinforcement learning
Original
Choi, J., Guo, Y., Moczulski, M., Oh, J., Wu, N., Norouzi, M., and Lee, H. (2018) · 2018
Later among the works it cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S. (2018) · 2018
Later among the works it cites.
Model-Based Reinforcement Learning via Meta-Policy Optimization
Clavera, I., Rothfuss, J., Schulman, J., Fujita, Y., Asfour, T., and Abbeel, P. (2018) · 2018
Later among the works it cites.
Efficient model-based deep reinforcement learning with variational state tabulation
Original
Corneil, D., Gerstner, W., and Brea, J. (2018) · 2018
Later among the works it cites.
End-to-end differentiable physics for learning and control
de Avila Belbute-Peres, F., Smith, K., Allen, K., Tenenbaum, J., and Kolter, J. Z. (2018) · 2018
Later among the works it cites.
Forward-backward reinforcement learning
Original
Edwards, A. D., Downs, L., and Davidson, J. C. (2018) · 2018
Later among the works it cites.
Beyond the One-Step Greedy Approach in Reinforcement Learning
Efroni, Y., Dalal, G., Scherrer, B., and Mannor, S. (2018) · 2018
Later among the works it cites.
Treeqn and atreec: Differentiable tree planning for deep reinforcement learning
Farquhar, G., Rocktäschel, T., Igl, M., and Whiteson, S. (2018) · 2018
Later among the works it cites.
Automatic Goal Generation for Reinforcement Learning Agents
Florensa, C., Held, D., Geng, X., and Abbeel, P. (2018) · 2018
Later among the works it cites.
Meta learning shared hierarchies
Frans, K., Ho, J., Chen, X., Abbeel, P., and Schulman, J. (2018) · 2018
Later among the works it cites.
Learning Actionable Representations with Goal-Conditioned Policies
Original
Ghosh, D., Gupta, A., and Levine, S. (2018) · 2018
Later among the works it cites.
Learning to search with MCTSnets
Original
Guez, A., Weber, T., Antonoglou, I., Simonyan, K., Vinyals, O., Wierstra, D., Munos, R., and Silver, D. (2018) · 2018
Later among the works it cites.
Recurrent world models facilitate policy evolution
Ha, D. and Schmidhuber, J. (2018) · 2018
Later among the works it cites.
Learning an Embedding Space for Transferable Robot Skills
Hausman, K., Springenberg, J. T., Wang, Z., Heess, N., and Riedmiller, M. (2018) · 2018
Later among the works it cites.
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D. (2018) · 2018
Later among the works it cites.
The effect of planning shape on dyna-style planning in high-dimensional state spaces
Original
Holland, G. Z., Talvitie, E. J., and Bowling, M. (2018) · 2018
Later among the works it cites.
Time-agnostic prediction: Predicting predictable video frames
Original
Jayaraman, D., Ebert, F., Efros, A. A., and Levine, S. (2018) · 2018
Later among the works it cites.
Feedback-Based Tree Search for Reinforcement Learning
Jiang, D., Ekwedike, E., and Liu, H. (2018) · 2018
Later among the works it cites.
Is Q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I. (2018) · 2018
Later among the works it cites.
Variance reduction methods for sublinear reinforcement learning
Kakade, S., Wang, M., and Yang, L. F. (2018) · 2018
Later among the works it cites.
Learning plannable representations with causal infogan
Kurutach, T., Tamar, A., Yang, G., Russell, S. J., and Abbeel, P. (2018) · 2018
Later among the works it cites.
Curiosity Driven Exploration of Learned Disentangled Goal Spaces
Laversanne-Finot, A., Pere, A., and Oudeyer, P.-Y. (2018) · 2018
Later among the works it cites.
State representation learning for control: An overview
Lesort, T., Díaz-Rodríguez, N., Goudou, J.-F., and Filliat, D. (2018) · 2018
Later among the works it cites.
Episodic memory deep Q-networks
Original
Lin, Z., Zhao, T., Yang, G., and Zhang, L. (2018) · 2018
Later among the works it cites.
Plan online, learn offline: Efficient learning and exploration via model-based control
Original
Lowrey, K., Rajeswaran, A., Kakade, S., Todorov, E., and Mordatch, I. (2018) · 2018
Later among the works it cites.
Now I Remember! Episodic Memory For Reinforcement Learning
Loynd, R., Hausknecht, M., Li, L., and Deng, L. (2018) · 2018
Later among the works it cites.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M. (2018) · 2018
Later among the works it cites.
Data-efficient hierarchical reinforcement learning
Nachum, O., Gu, S. S., Lee, H., and Levine, S. (2018) · 2018
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Nagabandi, A., Kahn, G., Fearing, R. S., and Levine, S. (2018c) · 2018
Later among the works it cites.
Adaptive skip intervals: Temporal abstraction for recurrent dynamical models
Neitz, A., Parascandolo, G., Bauer, S., and Schölkopf, B. (2018) · 2018
Later among the works it cites.
Unsupervised learning of goal spaces for intrinsically motivated goal exploration
Original
Péré, A., Forestier, S., Sigaud, O., and Oudeyer, P.-Y. (2018) · 2018
Later among the works it cites.
Temporal Difference Models: Model-Free Deep RL for Model-Based Control
Pong, V., Gu, S., Dalal, M., and Levine, S. (2018) · 2018
Later among the works it cites.
Learning abstract options
Riemer, M., Liu, M., and Tesauro, G. (2018) · 2018
Later among the works it cites.
Disentangling Controllable and Uncontrollable Factors of Variation by Interacting with the World
Original
Sawada, Y. (2018) · 2018
Later among the works it cites.
Time-contrastive networks: Self-supervised learning from video
Sermanet, P., Lynch, C., Chebotar, Y., Hsu, J., Jang, E., Schaal, S., Levine, S., and Brain, G. (2018) · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al. (2018) · 2018
Later among the works it cites.
Universal planning networks
Original
Srinivas, A., Jabri, A., Abbeel, P., Levine, S., and Finn, C. (2018) · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
Disentangling the independently controllable factors of variation by interacting with the world
Original
Thomas, V., Bengio, E., Fedus, W., Pondard, J., Beaudoin, P., Larochelle, H., Pineau, J., Precup, D., and Bengio, Y. (2018) · 2018
Later among the works it cites.
Contrastive explanations for reinforcement learning in terms of expected consequences
Original
van der Waa, J., van Diggelen, J., Bosch, K. v. d., and Neerincx, M. (2018) · 2018
Later among the works it cites.
Relational neural expectation maximization: Unsupervised discovery of objects and their interactions
Original
Van Steenkiste, S., Chang, M., Greff, K., and Schmidhuber, J. (2018) · 2018
Later among the works it cites.
Solving the Rubik’s cube with deep reinforcement learning and search
Agostinelli, F., McAleer, S., Shmakov, A., and Baldi, P. (2019) · 2019
Later among the works it cites.
Unsupervised state representation learning in atari
Anand, A., Racah, E., Ozair, S., Bengio, Y., Côté, M.-A., and Hjelm, R. D. (2019) · 2019
Later among the works it cites.
A differentiable physics engine for deep learning in robotics
Degrave, J., Hermans, M., Dambre, J., et al. (2019) · 2019
Later among the works it cites.
Feature control as intrinsic motivation for hierarchical reinforcement learning
Dilokthanakul, N., Kaplanis, C., Pawlowski, N., and Shanahan, M. (2019) · 2019
Later among the works it cites.
Combined reinforcement learning via abstract representations
François-Lavet, V., Bengio, Y., Precup, D., and Pineau, J. (2019) · 2019
Later among the works it cites.
An Investigation of Model-Free Planning
Guez, A., Mirza, M., Gregor, K., Kabra, R., Racaniere, S., Weber, T., Raposo, D., Santoro, A., Orseau, L., Eccles, T., et al. (2019) · 2019
Later among the works it cites.
Analogues of mental simulation and imagination in deep learning
Hamrick, J. B. (2019) · 2019
Later among the works it cites.
When to trust your model: Model-based policy optimization
Janner, M., Fu, J., Zhang, M., and Levine, S. (2019) · 2019
Later among the works it cites.
Hierarchical Reinforcement Learning with Hindsight
Levy, A., Platt, R., and Saenko, K. (2019) · 2019
Later among the works it cites.
Behaviour Suite for Reinforcement Learning
Osband, I., Doron, Y., Hessel, M., Aslanides, J., Sezener, E., Saraiva, A., McKinney, K., Lattimore, T., Szepezvari, C., Singh, S., Roy, B. V., Sutton, R., Silver, D., and Hasselt, H. V. (2019) · 2019
Later among the works it cites.
Dynamics-Aware Unsupervised Discovery of Skills
Sharma, A., Gu, S., Levine, S., Kumar, V., and Hausman, K. (2019) · 2019
Later among the works it cites.
Model-Based Active Exploration
Shyam, P., Jaśkowski, W., and Gomez, F. (2019) · 2019
Later among the works it cites.
When to use parametric models in reinforcement learning?
van Hasselt, H. P., Hessel, M., and Aslanides, J. (2019) · 2019
Later among the works it cites.
Meta-learning
Vanschoren, J. (2019) · 2019
Later among the works it cites.
Model-Based Multi-objective Reinforcement Learning with Unknown Weights
Yamaguchi, T., Nagahama, S., Ichikawa, Y., and Takadama, K. (2019) · 2019
Later among the works it cites.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Zanette, A. and Brunskill, E. (2019) · 2019
Later among the works it cites.
SOLAR: Deep Structured Representations for Model-Based Reinforcement Learning
Zhang, M., Vikram, S., Smith, L., Abbeel, P., Johnson, M., and Levine, S. (2019) · 2019
Later among the works it cites.
Amrl: Aggregated memory for reinforcement learning
Beck, J., Ciosek, K., Devlin, S., Tschiatschek, S., Zhang, C., and Hofmann, K. (2020) · 2020
Closest in time.
The Value Equivalence Principle for Model-Based Reinforcement Learning
Grimm, C., Barreto, A., Singh, S., and Silver, D. (2020) · 2020
Closest in time.
Combining q-learning and search with amortized value estimates
Hamrick, J. B., Bapst, V., Sanchez-Gonzalez, A., Pfaff, T., Weber, T., Buesing, L., and Battaglia, P. W. (2020) · 2020
Closest in time.
Contrastive Learning of Structured World Models
Kipf, T., van der Pol, E., and Welling, M. (2020) · 2020
Closest in time.
The LoCA Regret: A Consistent Metric to Evaluate Model-Based Behavior in Reinforcement Learning
Van Seijen, H., Nekoei, H., Racah, E., and Chandar, S. (2020) · 2020
Closest in time.
Value targets in off-policy AlphaZero: a new greedy backup
Willemsen, D., Baier, H., and Kaisers, M. (2020) · 2020
Closest in time.
Generalizable episodic memory for deep reinforcement learning
Original
Hu, H., Ye, J., Zhu, G., Ren, Z., and Zhang, C. (2021) · 2021
Closest in time.
World model as a graph: Learning latent landmarks for planning
Zhang, L., Yang, G., and Stadie, B. C. (2021) · 2021
Closest in time.