Fetching the paper…
Reading the bibliography…
Sample efficiency remains a fundamental issue of reinforcement learning.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M · 1912
Earlier work this paper cites.
Mixture density networks
Bishop, C. M · 1994
Earlier work this paper cites.
Mastering atari with discrete world models
Hafner, D., Lillicrap, T., Norouzi, M., and Ba, J · 2010
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2014
Earlier work this paper cites.
Convolutional lstm network: A machine learning approach for precipitation nowcasting
Shi, X., Chen, Z., Wang, H., Yeung, D.-Y., Wong, W.-k., and WOO, W.-c · 2015
Earlier work this paper cites.
Conditional image generation with pixelcnn decoders, 2016
van den Oord, A., Kalchbrenner, N., Vinyals, O., Espeholt, L., Graves, A., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Stochastic variational video prediction
Babaeizadeh, M., Finn, C., Erhan, D., Campbell, R. H., and Levine, S · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Neural discrete representation learning
van den Oord, A., Vinyals, O., and kavukcuoglu, k · 2017
Cited alongside, same era.
Predrnn: Recurrent neural networks for predictive learning using spatiotemporal lstms
Wang, Y., Long, M., Wang, J., Gao, Z., and Yu, P. S · 2017
Cited alongside, same era.
Recurrent world models facilitate policy evolution
Ha, D. and Schmidhuber, J · 2018
Later among the works it cites.
PredRNN++: Towards a resolution of the deep-in-time dilemma in spatiotemporal predictive learning
Wang, Y., Gao, Z., Long, M., Wang, J., and Philip, S. Y · 2018
Later among the works it cites.
The continuous bernoulli: fixing a pervasive error in variational autoencoders
Loaiza-Ganem, G. and Cunningham, J. P · 2019
Later among the works it cites.
Model based reinforcement learning for atari
Kaiser, L., Babaeizadeh, M., Miłos, P., Osiński, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., Mohiuddin, A., Sepassi, R., Tucker, G., and Michalewski, H · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…