Fetching the paper…
Reading the bibliography…
Planning has been very successful for control tasks with known environment dynamics.
Optimization of computer simulation models with rare events
Rubinstein, R. Y · 1997
Earlier work this paper cites.
Robust constrained model predictive control
Richards, A. G · 2005
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Nair, V. and Hinton, G. E · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Deisenroth, M. and Rasmussen, C. E · 2011
Earlier work this paper cites.
Synthesis and stabilization of complex behaviors through online trajectory optimization
Tassa, Y., Erez, T., and Todorov, E · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Cho, K., Van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, D. J., Mohamed, S., and Wierstra, D · 2014
Earlier work this paper cites.
Model regularization for stable sample rollouts
Talvitie, E · 2014
Earlier work this paper cites.
SOLAR: deep structured representations for model-based reinforcement learning
Zhang, M., Vikram, S., Smith, L., Abbeel, P., Johnson, M., and Levine, S · 2014
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Bengio, S., Vinyals, O., Jaitly, N., and Shazeer, N · 2015
Earlier work this paper cites.
A recurrent latent variable model for sequential data
Chung, J., Kastner, K., Dinh, L., Goel, K., Courville, A. C., and Bengio, Y · 2015
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Clevert, D.-A., Unterthiner, T., and Hochreiter, S · 2015
Earlier work this paper cites.
Krishnan, R. G., Shalit, U., and Sontag, D · 2015
Earlier work this paper cites.
Deep multi-scale video prediction beyond mean square error
Mathieu, M., Couprie, C., and LeCun, Y · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
Oh, J., Guo, X., Lee, H., Lewis, R. L., and Singh, S · 2015
Earlier work this paper cites.
Improving multi-step prediction of learned time series models
Venkatraman, A., Hebert, M., and Bagnell, J. A · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Watter, M., Springenberg, J., Boedecker, J., and Riedmiller, M · 2015
Earlier work this paper cites.
Learning to poke by poking: Experiential learning of intuitive physics
Agrawal, P., Nair, A. V., Abbeel, P., Malik, J., and Levine, S · 2016
Cited alongside, same era.
Improving pilco with bayesian neural network dynamics models
Gal, Y., McAllister, R., and Rasmussen, C. E · 2016
Cited alongside, same era.
beta-vae: Learning basic visual concepts with a constrained variational framework
Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., and Lerchner, A · 2016
Cited alongside, same era.
Kalchbrenner, N., Oord, A. v. d., Simonyan, K., Danihelka, I., Vinyals, O., Graves, A., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Deep variational bayes filters: Unsupervised learning of state space models from raw data
Karl, M., Soelch, M., Bayer, J., and van der Smagt, P · 2016
Cited alongside, same era.
Neural discrete representation learning
van den Oord, A., Vinyals, O., et al · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Weber, T., Racanière, S., Reichert, D. P., Buesing, L., Guez, A., Rezende, D. J., Badia, A. P., Vinyals, O., Heess, N., Li, Y., et al · 2017
Later among the works it cites.
Learning awareness models
Amos, B., Dinh, L., Cabi, S., Rothörl, T., Muldal, A., Erez, T., Tassa, Y., de Freitas, N., and Denil, M · 2018
Closest in time.
Distributed distributional deterministic policy gradients
Barth-Maron, G., Hoffman, M. W., Budden, D., Dabney, W., Horgan, D., Muldal, A., Heess, N., and Lillicrap, T · 2018
Closest in time.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Buckman, J., Hafner, D., Tucker, G., Brevdo, E., and Lee, H · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Professor forcing: A new algorithm for training recurrent networks
Lamb, A. M., GOYAL, A. G. A. P., Zhang, Y., Zhang, S., Courville, A. C., and Bengio, Y · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Generating videos with scene dynamics
Vondrick, C., Pirsiavash, H., and Torralba, A · 2016
Cited alongside, same era.
Stochastic variational video prediction
Babaeizadeh, M., Finn, C., Erhan, D., Campbell, R. H., and Levine, S · 2017
Cited alongside, same era.
Robust locally-linear controllable embedding
Banijamali, E., Shu, R., Ghavamzadeh, M., Bui, H., and Ghodsi, A · 2017
Cited alongside, same era.
Recurrent environment simulators
Chiappa, S., Racaniere, S., Wierstra, D., and Mohamed, S · 2017
Cited alongside, same era.
Dillon, J. V., Langmore, I., Tran, D., Brevdo, E., Vasudevan, S., Moore, D., Patton, B., Alemi, A., Hoffman, M., and Saurous, R. A · 2017
Cited alongside, same era.
Closest in time.
Learning and querying fast generative models for reinforcement learning
Buesing, L., Weber, T., Racaniere, S., Eslami, S., Rezende, D., Reichert, D. P., Viola, F., Besse, F., Gregor, K., Hassabis, D., et al · 2018
Closest in time.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S · 2018
Closest in time.
Stochastic video generation with a learned prior
Denton, E. and Fergus, R · 2018
Closest in time.
Probabilistic recurrent state-space models
Doerr, A., Daniel, C., Schiegg, M., Nguyen-Tuong, D., Schaal, S., Toussaint, M., and Trimpe, S · 2018
Closest in time.
Visual foresight: Model-based deep reinforcement learning for vision-based robotic control
Ebert, F., Finn, C., Dasari, S., Xie, A., Lee, A., and Levine, S · 2018
Closest in time.
Temporal difference variational auto-encoder
Gregor, K. and Besse, F · 2018
Closest in time.
Ha, D. and Schmidhuber, J · 2018
Closest in time.
Model-based planning with discrete and continuous actions
Henaff, M., Whitney, W. F., and LeCun, Y · 2018
Closest in time.
Synthesizing neural network controllers with probabilistic model based reinforcement learning
Higuera, J. C. G., Meger, D., and Dudek, G · 2018
Closest in time.
Deep variational reinforcement learning for pomdps
Igl, M., Zintgraf, L., Le, T. A., Wood, F., and Whiteson, S · 2018
Closest in time.
Glow: Generative flow with invertible 1x1 convolutions
Kingma, D. P. and Dhariwal, P · 2018
Closest in time.
Model-ensemble trust-region policy optimization
Kurutach, T., Clavera, I., Duan, Y., Tamar, A., and Abbeel, P · 2018
Closest in time.
Srinivas, A., Jabri, A., Abbeel, P., Levine, S., and Finn, C · 2018
Closest in time.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al · 2018
Closest in time.
Unsupervised predictive memory in a goal-directed agent
Wayne, G., Hung, C.-C., Amos, D., Mirza, M., Ahuja, A., Grabska-Barwinska, A., Rae, J., Mirowski, P., Leibo, J. Z., Santoro, A., et al · 2018
Closest in time.