Fetching the paper…
Reading the bibliography…
In recent years there is a growing interest in using deep representations for reinforcement learning.
Some methods for classification and analysis of multivariate observations
James MacQueen et al · 1967
Earlier work this paper cites.
The internal model principle for linear multivariable regulators
Bruce A Francis and William M Wonham · 1975
Earlier work this paper cites.
Variable resolution dynamic programming: Efficiently learning action maps in multivariate real-valued state-spaces
Andrew Moore · 1991
Earlier work this paper cites.
Feudal reinforcement learning
Peter Dayan and Geoffrey E Hinton · 1993
Earlier work this paper cites.
Reinforcement learning for robots using neural networks
Long-Ji Lin · 1993
Earlier work this paper cites.
Decomposition techniques for planning in stochastic domains
Thomas Dean and Shieu-Hong Lin · 1995
Earlier work this paper cites.
Stable function approximation in dynamic programming
Geoffrey J Gordon · 1995
Earlier work this paper cites.
Reinforcement learning with soft state aggregation
Satinder P Singh, Tommi Jaakkola, and Michael I Jordan · 1995
Earlier work this paper cites.
Temporal difference learning and TD-Gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
John N Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
Hierarchical solution of Markov decision processes using macro-actions
Milos Hauskrecht, Nicolas Meuleau, Leslie Pack Kaelbling, Thomas Dean, and Craig Boutilier · 1998
Earlier work this paper cites.
Flexible decomposition algorithms for weakly coupled Markov decision problems
Ronald Parr · 1998
Earlier work this paper cites.
Learning metric-topological maps for indoor mobile robot navigation
Sebastian Thrun · 1998
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the MAXQ value function decomposition
Thomas G Dietterich · 2000
Earlier work this paper cites.
A global geometric framework for nonlinear dimensionality reduction
Joshua B Tenenbaum, Vin De Silva, and John C Langford · 2000
Earlier work this paper cites.
Learning embedded maps of markov processes
Yaakov Engel and Shie Mannor · 2001
Earlier work this paper cites.
Q-cut—dynamic discovery of sub-goals in reinforcement learning
Ishai Menache, Shie Mannor, and Nahum Shimkin · 2002
Earlier work this paper cites.
Adaptation and regulation with signal detection implies internal model
Eduardo D Sontag · 2003
Cited alongside, same era.
Dynamic abstraction in reinforcement learning via clustering
Shie Mannor, Ishai Menache, Amit Hoze, and Uri Klein · 2004
Cited alongside, same era.
Automated discovery of options in reinforcement learning
Martin Stolle · 2004
Cited alongside, same era.
Invariant visual representation by single neurons in the human brain
R Quian Quiroga, Leila Reddy, Gabriel Kreiman, Christof Koch, and Itzhak Fried · 2005
Cited alongside, same era.
Neural fitted Q iteration–first experiences with a data efficient neural reinforcement learning method
Martin Riedmiller · 2005
Cited alongside, same era.
Identifying useful subgoals in reinforcement learning by local graph partitioning
Özgür Şimşek, Alicia P Wolfe, and Andrew G Barto · 2005
Accelerating t-SNE using tree-based algorithms
Laurens Van Der Maaten · 2014
Later among the works it cites.
How transferable are features in deep neural networks?
Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson · 2014
Later among the works it cites.
Visualizing and understanding convolutional networks
Matthew D Zeiler and Rob Fergus · 2014
Later among the works it cites.
Increasing the action gap: New operators for reinforcement learning
Marc G Bellemare, Georg Ostrovski, Arthur Guez, Philip S Thomas, and Rémi Munos · 2015
Later among the works it cites.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A tutorial on spectral clustering
Ulrike Von Luxburg · 2007
Cited alongside, same era.
Visualizing higher-layer features of a deep network
Dumitru Erhan, Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2009
Cited alongside, same era.
Towards perceptual shared autonomy for robotic mobile manipulation
Benjamin Pitzer, Michael Styer, Christian Bersch, Charles DuHadway, and Jan Becker · 2011
Cited alongside, same era.
Mayavi: 3D Visualization of Scientific Data
P. Ramachandran and G. Varoquaux · 2011
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2012
Cited alongside, same era.
Model selection in markovian processes
Assaf Hallak, Dotan Di-Castro, and Shie Mannor · 2013
Cited alongside, same era.
Approximate value iteration with temporally extended actions
Timothy A Mann, Shie Mannor, and Doina Precup · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Later among the works it cites.
Massively parallel methods for deep reinforcement learning
Arun Nair, Praveen Srinivasan, Sam Blackwell, Cagdas Alcicek, Rory Fearon, Alessandro De Maria, Vedavyas Panneershelvam, Mustafa Suleyman, Charles Beattie, Stig Petersen, et al · 2015
Later among the works it cites.
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2015
Later among the works it cites.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2015
Later among the works it cites.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Nando de Freitas, and Marc Lanctot · 2015
Later among the works it cites.
Spatio-temporal abstractions in reinforcement learning through neural encoding
Nir Baram, Tom Zahavy, and Shie Mannor · 2016
Closest in time.
Visualizing dynamics: from t-sne to semi-mdps
Nir Ben Zrihem, Tom Zahavy, and Shie Mannor · 2016
Closest in time.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas D Kulkarni, Karthik R Narasimhan, Ardavan Saeedi, and Joshua B Tenenbaum · 2016
Closest in time.
Adaptive skills, adaptive partitions (asap)
Daniel J Mankowitz, Timothy A Mann, and Shie Mannor · 2016
Closest in time.
Graying the black box: Understanding dqns
Tom Zahavy, Nir Ben Zrihem, and Shie Mannor · 2016
Closest in time.
A deep hierarchical approach to lifelong learning in minecraft
Chen Tessler, Shahar Givony, Tom Zahavy, Daniel J Mankowitz, and Shie Mannor · 2017
Closest in time.