Fetching the paper…
Reading the bibliography…
Model-based Reinforcement Learning approaches have the promise of being sample efficient.
Dyna, an integrated architecture for learning, planning, and reacting
Sutton, R. S · 1991
Earlier work this paper cites.
On-line Q-learning using connectionist systems , volume 37
Rummery, G. A. and Niranjan, M · 1994
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Sample-based learning and search with permanent and transient memories
Silver, D., Sutton, R. S., and Müller, M · 2008
Earlier work this paper cites.
Where science starts: Spontaneous experiments in preschoolers’ exploratory play
Cook, C., Goodman, N. D., and Schulz, L. E · 2011
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Deisenroth, M. and Rasmussen, C. E · 2011
Earlier work this paper cites.
Dyna-style planning with linear function approximation and prioritized sweeping
Sutton, R. S., Szepesvári, C., Geramifard, A., and Bowling, M · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Reinforcement learning in robust markov decision processes
Lim, S. H., Xu, H., and Mannor, S · 2013
Earlier work this paper cites.
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
Cho, K., van Merrienboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Earlier work this paper cites.
Model regularization for stable sample rollouts
Talvitie, E · 2014
Cited alongside, same era.
Scheduled sampling for sequence prediction with recurrent neural networks
Bengio, S., Vinyals, O., Jaitly, N., and Shazeer, N · 2015
Cited alongside, same era.
Automated adaptive inference of phenomenological dynamical models
Daniels, B. C. and Nemenman, I · 2015
Cited alongside, same era.
Deepmpc: Learning deep latent features for model predictive control
Lenz, I., Knepper, R. A., and Saxena, A · 2015
Cited alongside, same era.
Ensemble-cio: Full-body dynamic motion planning that transfers to physical humanoids
Mordatch, I., Lowrey, K., and Todorov, E · 2015
Cited alongside, same era.
Openai gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Epopt: Learning robust neural network policies using model ensembles
Rajeswaran, A., Ghotra, S., Ravindran, B., and Levine, S · 2016
Later among the works it cites.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Pieter Abbeel, O., and Zaremba, W · 2017
Later among the works it cites.
Openai baselines
Dhariwal, P., Hesse, C., Klimov, O., Nichol, A., Plappert, M., Radford, A., Schulman, J., Sidor, S., Wu, Y., and Zhokhov, P · 2017
Later among the works it cites.
Uncertainty-driven imagination for continuous deep reinforcement learning
Kalweit, G. and Boedecker, J · 2017
Later among the works it cites.
Neural Network Dynamics for Model-Based Deep Reinforcement Learning with Model-Free Fine-Tuning
Nagabandi, A., Kahn, G., Fearing, R. S., and Levine, S · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Benchmarking Deep Reinforcement Learning for Continuous Control
Duan, Y., Chen, X., Houthooft, R., Schulman, J., and Abbeel, P · 2016
Cited alongside, same era.
Reinforcement Learning with Unsupervised Auxiliary Tasks
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Professor Forcing: A New Algorithm for Training Recurrent Networks
Lamb, A., Goyal, A., Zhang, Y., Zhang, S., Courville, A., and Bengio, Y · 2016
Cited alongside, same era.
Asynchronous Methods for Deep Reinforcement Learning
Mnih, V., Puigdomènech Badia, A., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al
Cited in the paper.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al
Cited in the paper.
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Weber, T., Racanière, S., Reichert, D. P., Buesing, L., Guez, A., Rezende, D. J., Badia, A. P., Vinyals, O., Heess, N., Li, Y., et al · 2017
Later among the works it cites.
Learning and Querying Fast Generative Models for Reinforcement Learning
Buesing, L., Weber, T., Racaniere, S., Eslami, S. M. A., Rezende, D., Reichert, D. P., Viola, F., Besse, F., Gregor, K., Hassabis, D., and Wierstra, D · 2018
Later among the works it cites.
Multi-goal reinforcement learning: Challenging robotics environments and request for research, 2018
Plappert, M., Andrychowicz, M., Ray, A., McGrew, B., Baker, B., Powell, G., Schneider, J., Tobin, J., Chociej, M., Welinder, P., Kumar, V., and Zaremba, W · 2018
Later among the works it cites.