Fetching the paper…
Reading the bibliography…
In this paper we study how to learn stochastic, multimodal transition dynamics in reinforcement learning (RL) tasks.
Sutton, R.S.: Dyna, an integrated architecture for learning, planning, and reacting. ACM SIGART Bulletin 2(4), 160–163 (1991)
1991
Earlier work this paper cites.
Atkeson, C.G., Moore, A.W., Schaal, S.: Locally weighted learning for control. In: Lazy learning, pp. 75–113. Springer (1997)
1997
Earlier work this paper cites.
Brafman, R.I., Tennenholtz, M.: R-max-a general polynomial time algorithm for near-optimal reinforcement learning. Journal of Machine Learning Research 3(Oct), 213–231 (2002)
2002
Earlier work this paper cites.
Deisenroth, M., Rasmussen, C.E.: PILCO: A model-based and data-efficient approach to policy search. In: Proceedings of the 28th International Conference on machine learning (ICML-11). pp. 465–472 (2011)
2011
Earlier work this paper cites.
Li, L., Littman, M.L., Walsh, T.J., Strehl, A.L.: Knows what it knows: a framework for self-aware learning. Machine learning 82(3), 399–443 (2011)
2011
Earlier work this paper cites.
Hester, T., Stone, P.: Learning and using models. In: Reinforcement Learning, pp. 111–141. Springer (2012)
2012
Earlier work this paper cites.
Kingma, D.P., Welling, M.: Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114 (2013)
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
Chung, J., Kastner, K., Dinh, L., Goel, K., Courville, A.C., Bengio, Y.: A recurrent latent variable model for sequential data. In: Advances in neural information processing systems. pp. 2980–2988 (2015)
2015
Earlier work this paper cites.
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A.A., Veness, J., Bellemare, M.G., Graves, A., Riedmiller, M., Fidjeland, A.K., Ostrovski, G., et al.: Human-level control through deep reinforcement learning. Nature 518(7540), 529–533 (2015)
2015
Earlier work this paper cites.
Oh, J., Guo, X., Lee, H., Lewis, R.L., Singh, S.: Action-conditional video prediction using deep networks in atari games. In: Advances in Neural Information Processing Systems. pp. 2863–2871 (2015)
2015
Cited alongside, same era.
2015
Cited alongside, same era.
Sohn, K., Lee, H., Yan, X.: Learning structured output representation using deep conditional generative models. In: Advances in Neural Information Processing Systems. pp. 3483–3491 (2015)
2015
Cited alongside, same era.
2015
Cited alongside, same era.
Houthooft, R., Chen, X., Duan, Y., Schulman, J., De Turck, F., Abbeel, P.: VIME: Variational information maximizing exploration. In: Advances in Neural Information Processing Systems. pp. 1109–1117 (2016)
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
Kingma, D.P., Salimans, T., Jozefowicz, R., Chen, X., Sutskever, I., Welling, M.: Improved Variational Inference with Inverse Autoregressive Flow. In: Advances in Neural Information Processing Systems. pp. 4743–4751 (2016)
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Watter, M., Springenberg, J., Boedecker, J., Riedmiller, M.: Embed to control: A locally linear latent dynamics model for control from raw images. In: Advances in Neural Information Processing Systems. pp. 2746–2754 (2015)
2015
Cited alongside, same era.
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., Munos, R.: Unifying count-based exploration and intrinsic motivation. In: Advances in Neural Information Processing Systems. pp. 1471–1479 (2016)
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
Gal, Y., McAllister, R.T., Rasmussen, C.E.: Improving PILCO with Bayesian Neural Network Dynamics Models. In: Data-Efficient Machine Learning workshop. vol. 951, p. 2016 (2016)
2016
Cited alongside, same era.
2016
Cited alongside, same era.
Li, Y., Turner, R.E.: Rényi divergence variational inference. In: Advances in Neural Information Processing Systems. pp. 1073–1081 (2016)
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
Sønderby, C.K., Raiko, T., Maaløe, L., Sønderby, S.K., Winther, O.: Ladder variational autoencoders. In: Advances in Neural Information Processing Systems. pp. 3738–3746 (2016)
2016
Later among the works it cites.
Walker, J., Doersch, C., Gupta, A., Hebert, M.: An Uncertain Future: Forecasting from Static Images Using Variational Autoencoders. In: European Conference on Computer Vision. pp. 835–851. Springer (2016)
2016
Later among the works it cites.
Finn, C., Levine, S.: Deep visual foresight for planning robot motion. In: IEEE International Conference on Robotics and Automation (ICRA) (2017)
2017
Closest in time.