Fetching the paper…
Reading the bibliography…
We present an approach to sensorimotor control in immersive environments.
Learning to predict by the methods of temporal differences
Richard S. Sutton · 1988
Earlier work this paper cites.
Forward models: Supervised learning with a distal teacher
Michael I. Jordan and David E. Rumelhart · 1992
Earlier work this paper cites.
TD-gammon, a self-teaching backgammon program, achieves master-level play
Gerald Tesauro · 1994
Earlier work this paper cites.
A counterexample to temporal differences learning
Dimitri P. Bertsekas · 1995
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding
Richard S. Sutton · 1995
Earlier work this paper cites.
Reinforcement learning: A survey
Leslie Pack Kaelbling, Michael L. Littman, and Andrew W. Moore · 1996
Earlier work this paper cites.
Reinforcement learning with replacing eligibility traces
Satinder P. Singh and Richard S. Sutton · 1996
Earlier work this paper cites.
Multi-criteria reinforcement learning
Zoltán Gábor, Zsolt Kalmár, and Csaba Szepesvári · 1998
Earlier work this paper cites.
A unified analysis of value-function-based reinforcement learning algorithms
Csaba Szepesvári and Michael L. Littman · 1999
Earlier work this paper cites.
Predictive representations of state
Michael L. Littman, Richard S. Sutton, and Satinder P. Singh · 2001
Earlier work this paper cites.
On the convergence of optimistic policy iteration
John N. Tsitsiklis · 2002
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
Andrew G. Barto and Sridhar Mahadevan · 2003
Earlier work this paper cites.
Learning rates for Q-learning
Eyal Even-Dar and Yishay Mansour · 2003
Earlier work this paper cites.
Masters of Doom: How Two Guys Created an Empire and Transformed Pop Culture
David Kushner · 2003
Earlier work this paper cites.
Learning predictive state representations
Satinder P. Singh, Michael L. Littman, Nicholas K. Jong, David Pardoe, and Peter Stone · 2003
Earlier work this paper cites.
Off-road obstacle avoidance through end-to-end learning
Yann LeCun, Urs Muller, Jan Ben, Eric Cosatto, and Beat Flepp · 2005
Earlier work this paper cites.
Motor development
Karen E. Adolph and Sarah E. Berger · 2006
Cited alongside, same era.
Pathologies of temporal difference methods in approximate dynamic programming
Dimitri P. Bertsekas · 2010
Cited alongside, same era.
Horde: a scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Richard S. Sutton, Joseph Modayil, Michael Delp, Thomas Degris, Patrick M. Pilarski, Adam White, and Doina Precup · 2011
Cited alongside, same era.
Learning parameterized skills
Bruno Castro da Silva, George Konidaris, and Andrew G. Barto · 2012
Cited alongside, same era.
Reinforcement learning to adjust parametrized motor primitives to new situations
Jens Kober, Andreas Wilhelm, Erhan Oztop, and Jan Peters · 2012
Cited alongside, same era.
Transfer in reinforcement learning via shared features
George Konidaris, Ilya Scheidwasser, and Andrew G. Barto · 2012
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Later among the works it cites.
Charles Blundell, Benigno Uria, Alexander Pritzel, Yazhe Li, Avraham Ruderman, Joel Z. Leibo, Jack Rae, Daan Wierstra, and Demis Hassabis · 2016
Closest in time.
Unsupervised learning for physical interaction through video prediction
Chelsea Finn, Ian J. Goodfellow, and Sergey Levine · 2016
Closest in time.
Nal Kalchbrenner, Aaron van den Oord, Karen Simonyan, Ivo Danihelka, Oriol Vinyals, Alex Graves, and Koray Kavukcuoglu · 2016
Closest in time.
ViZDoom: A Doom-based AI research platform for visual reinforcement learning
Michał Kempka, Marek Wydmuch, Grzegorz Runc, Jakub Toczek, and Wojciech Jaśkowski · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Machine Learning: A Probabilistic Perspective
Kevin P. Murphy · 2012
Cited alongside, same era.
Reinforcement learning in robotics: A survey
Jens Kober, J. Andrew Bagnell, and Jan Peters · 2013
Cited alongside, same era.
Guided policy search
Sergey Levine and Vladlen Koltun · 2013
Cited alongside, same era.
A survey of multi-objective sequential decision-making
Diederik M. Roijers, Peter Vamplew, Shimon Whiteson, and Richard Dazeley · 2013
Cited alongside, same era.
Learning monocular reactive UAV control in cluttered natural environments
Stéphane Ross, Narek Melik-Barkhudarov, Kumar Shaurya Shankar, Andreas Wendel, Debadeepta Dey, J. Andrew Bagnell, and Martial Hebert · 2013
Cited alongside, same era.
Multi-task policy search for robotics
Marc Peter Deisenroth, Peter Englert, Jan Peters, and Dieter Fox · 2014
Cited alongside, same era.
Brenden M. Lake, Tomer D. Ullman, Joshua B. Tenenbaum, and Samuel J. Gershman · 2016
Closest in time.
Playing FPS games with deep reinforcement learning
Guillaume Lample and Devendra Singh Chaplot · 2016
Closest in time.
Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection
Sergey Levine, Peter Pastor, Alex Krizhevsky, and Deirdre Quillen · 2016
Closest in time.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Closest in time.
Deep multi-scale video prediction beyond mean square error
Michaël Mathieu, Camille Couprie, and Yann LeCun · 2016
Closest in time.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Closest in time.
Control of memory, active perception, and action in Minecraft
Junhyuk Oh, Valliappa Chockalingam, Satinder P. Singh, and Honglak Lee · 2016
Closest in time.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, et al · 2016
Closest in time.
WaveNet: A generative model for raw audio
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu · 2016
Closest in time.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot, and Nando de Freitas · 2016
Closest in time.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2017
Closest in time.