Fetching the paper…
Reading the bibliography…
We introduce a method which enables a recurrent dynamics model to be temporally abstract.
Markov decision processes. j
Puterman, M. L. (1994) · 1994
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S. (1999) · 1999
Earlier work this paper cites.
Frame skip is a powerful parameter for learning to play atari
Braylan, A., Hollenbeck, M., Meyerson, E., and Miikkulainen, R. (2000) · 2000
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
Barto, A. G. and Mahadevan, S. (2003) · 2003
Earlier work this paper cites.
Universal intelligence: A definition of machine intelligence
Legg, S. and Hutter, M. (2007) · 2007
Earlier work this paper cites.
Dynamic time warping
Müller, M. (2007) · 2007
Earlier work this paper cites.
Causality
Pearl, J. (2009) · 2009
Earlier work this paper cites.
Deep sparse rectifier neural networks
Glorot, X., Bordes, A., and Bengio, Y. (2011) · 2011
Earlier work this paper cites.
Model-based reinforcement learning as cognitive search: neurocomputational theories
Daw, N. D. (2012) · 2012
Earlier work this paper cites.
On causal and anticausal learning
Schölkopf, B., Janzing, D., Peters, J., Sgouritsa, E., Zhang, K., and Mooij, J. (2012) · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Bengio, S., Vinyals, O., Jaitly, N., and Shazeer, N. (2015) · 2015
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J. (2015) · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in atari games
Oh, J., Guo, X., Lee, H., Lewis, R. L., and Singh, S. (2015) · 2015
Cited alongside, same era.
Embed to control: A locally linear latent dynamics model for control from raw images
Watter, M., Springenberg, J., Boedecker, J., and Riedmiller, M. (2015) · 2015
Cited alongside, same era.
Time-adaptive cross entropy planning
Belzner, L. (2016) · 2016
Cited alongside, same era.
Deep visual foresight for planning robot motion
Finn, C. and Levine, S. (2017) · 2017
Later among the works it cites.
Sparse attentive backtracking: Long-range credit assignment in recurrent networks
Ke, N. R., Goyal, A., Bilaniuk, O., Binas, J., Charlin, L., Pal, C., and Bengio, Y. (2017) · 2017
Later among the works it cites.
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M. (2017) · 2017
Later among the works it cites.
Value prediction network
Oh, J., Singh, S., and Lee, H. (2017) · 2017
Later among the works it cites.
Learning independent causal mechanisms
Parascandolo, G., Rojas-Carulla, M., Kilbertus, N., and Schölkopf, B. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Heess, N., Wayne, G., Tassa, Y., Lillicrap, T., Riedmiller, M., and Silver, D. (2016) · 2016
Cited alongside, same era.
Causal inference by using invariant prediction: identification and confidence intervals
Peters, J., Bühlmann, P., and Meinshausen, N. (2016) · 2016
Cited alongside, same era.
The predictron: End-to-end learning and planning
Silver, D., van Hasselt, H., Hessel, M., Schaul, T., Guez, A., Harley, T., Dulac-Arnold, G., Reichert, D., Rabinowitz, N., Barreto, A., et al. (2016) · 2016
Cited alongside, same era.
A brief survey of deep reinforcement learning
Arulkumaran, K., Deisenroth, M. P., Brundage, M., and Bharath, A. A. (2017) · 2017
Cited alongside, same era.
Recurrent environment simulators
Chiappa, S., Racaniere, S., Wierstra, D., and Mohamed, S. (2017) · 2017
Cited alongside, same era.
Self-supervised visual planning with temporal skip connections
Ebert, F., Finn, C., Lee, A. X., and Levine, S. (2017) · 2017
Cited alongside, same era.
Elements of causal inference: foundations and learning algorithms
Peters, J., Janzing, D., and Schölkopf, B. (2017) · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Weber, T., Racanière, S., Reichert, D. P., Buesing, L., Guez, A., Rezende, D. J., Badia, A. P., Vinyals, O., Heess, N., Li, Y., et al. (2017) · 2017
Later among the works it cites.
Learning and querying fast generative models for reinforcement learning
Buesing, L., Weber, T., Racaniere, S., Eslami, S., Rezende, D., Reichert, D. P., Viola, F., Besse, F., Gregor, K., Hassabis, D., et al. (2018) · 2018
Closest in time.
Time-agnostic prediction: Predicting predictable video frames
Jayaraman, D., Ebert, F., Efros, A. A., and Levine, S. (2018) · 2018
Closest in time.
Temporal difference models: Model-free deep rl for model-based control
Pong, V., Gu, S., Dalal, M., and Levine, S. (2018) · 2018
Closest in time.