Fetching the paper…
Reading the bibliography…
One of the key challenges of artificial intelligence is to learn models that are effective in the context of planning.
Dynamic programming
Bellman, Richard · 1957
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Integrated architectures for learning, planning and reacting based on dynamic programming
Sutton, R. S · 1990
Earlier work this paper cites.
TD models: Modeling the world at a mixture of time scales
Sutton, Richard S · 1995
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Yann, Bottou, Léon, Bengio, Yoshua, and Haffner, Patrick · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Predictive representations of state
Littman, Michael L, Sutton, Richard S, and Singh, Satinder P · 2001
Earlier work this paper cites.
Deep sparse rectifier neural networks
Glorot, Xavier, Bordes, Antoine, and Bengio, Yoshua · 2011
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, Richard S, Modayil, Joseph, Delp, Michael, Degris, Thomas, Pilarski, Patrick M, White, Adam, and Precup, Doina · 2011
Earlier work this paper cites.
Multi-timescale nexting in a reinforcement learning robot
Modayil, Joseph, White, Adam, and Sutton, Richard S · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, Emanuel, Erez, Tom, and Tassa, Yuval · 2012
Cited alongside, same era.
Better Generalization with Forecasts
Schaul, Tom and Ring, Mark B · 2013
Cited alongside, same era.
Deep residual learning for image recognition
He, Kaiming, Zhang, Xiangyu, Ren, Shaoqing, and Sun, Jian · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, Sergey and Szegedy, Christian · 2015
Cited alongside, same era.
A method for stochastic optimization
Action-conditional video prediction using deep networks in atari games
Oh, Junhyuk, Guo, Xiaoxiao, Lee, Honglak, Lewis, Richard L, and Singh, Satinder · 2015
Later among the works it cites.
Universal Value Function Approximators
Schaul, Tom, Horgan, Daniel, Gregor, Karol, and Silver, David · 2015
Later among the works it cites.
Schmidhuber, Juergen · 2015
Later among the works it cites.
Recurrent environment simulators
Chiappa, Silvia, Racaniere, Sebastien, Wierstra, Daan, and Mohamed, Shakir · 2016
Closest in time.
Adaptive computation time for recurrent neural networks
Graves, Alex · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kingma, Diederik P and Ba, Jimmy · 2015
Cited alongside, same era.
Deeply-supervised nets
Lee, Chen-Yu, Xie, Saining, Gallagher, Patrick, Zhang, Zhengyou, and Tu, Zhuowen · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A., Veness, Joel, Bellemare, Marc G., Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K., Ostrovski, Georg, Petersen, Stig, Beattie, Charles, Sadik, Amir, Antonoglou, Ioannis, King, Helen, Kumaran, Dharshan, Wierstra, Daan, Legg, Shane, and Hassabis, Demis · 2015
Cited alongside, same era.
Lillicrap, T., Hunt, J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Closest in time.
Asynchronous methods for deep reinforcement learning
Mnih, V, Badia, A Puigdomènech, Mirza, M, Graves, A, Lillicrap, T, Harley, T, Silver, D, and Kavukcuoglu, K · 2016
Closest in time.
Value iteration networks
Tamar, Aviv, Wu, Yi, Thomas, Garrett, Levine, Sergey, and Abbeel, Pieter · 2016
Closest in time.