Fetching the paper…
Reading the bibliography…
Goals for reinforcement learning problems are typically defined through hand-specified rewards.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, Richard S · 1990
Earlier work this paper cites.
Prioritized sweeping: Reinforcement learning with less data and less time
Moore, Andrew W and Atkeson, Christopher G · 1993
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Sutton, Richard S and Barto, Andrew G · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, Andrew Y, Harada, Daishi, and Russell, Stuart · 1999
Earlier work this paper cites.
Forward and bidirectional planning based on reinforcement learning and neural networks in a simulated robot
Baldassarre, Gianluca · 2003
Earlier work this paper cites.
Horizon-based value iteration
Zang, Peng, Irani, Arya, and Isbell Jr, Charles Lee · 2007
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, et al · 2015
Cited alongside, same era.
Schaul, Tom, Quan, John, Antonoglou, Ioannis, and Silver, David · 2015
Cited alongside, same era.
Learning to poke by poking: Experiential learning of intuitive physics
Agrawal, Pulkit, Nair, Ashvin V, Abbeel, Pieter, Malik, Jitendra, and Levine, Sergey · 2016
Cited alongside, same era.
Continuous deep q-learning with model-based acceleration
Gu, Shixiang, Lillicrap, Timothy, Sutskever, Ilya, and Levine, Sergey · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Van Hasselt, Hado, Guez, Arthur, and Silver, David · 2016
Later among the works it cites.
Reverse curriculum generation for reinforcement learning
Florensa, Carlos, Held, David, Wulfmeier, Markus, and Abbeel, Pieter · 2017
Later among the works it cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, Deepak, Agrawal, Pulkit, Efros, Alexei A, and Darrell, Trevor · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Weber, Théophane, Racanière, Sébastien, Reichert, David P, Buesing, Lars, Guez, Arthur, Rezende, Danilo Jimenez, Badia, Adria Puigdomènech, Vinyals, Oriol, Heess, Nicolas, Li, Yujia, et al · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…