Fetching the paper…
Reading the bibliography…
When function approximation is deployed in reinforcement learning (RL), the same problem may be formulated in different ways, often by treating a pre-processing step as a part of the environment or as part of the agent.
Decision theoretic generalizations of the PAC model for neural net and other learning applications
David Haussler · 1992
Earlier work this paper cites.
Markov Decision Processes
ML Puterman · 1994
Earlier work this paper cites.
Feature-based methods for large scale dynamic programming
Benjamin Van Roy · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
Stable function approximation in dynamic programming
Geoffrey J Gordon · 1995
Earlier work this paper cites.
The lack of a priori distinctions between learning algorithms
David H Wolpert · 1996
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
John N Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
PEGASUS: A policy search method for large MDPs and POMDPs
Andrew Y Ng and Michael Jordan · 2000
Earlier work this paper cites.
Approximately Optimal Approximate Reinforcement Learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
A sparse sampling algorithm for near-optimal planning in large Markov decision processes
Michael Kearns, Yishay Mansour, and Andrew Y Ng · 2002
Earlier work this paper cites.
Error bounds for approximate policy iteration
Rémi Munos · 2003
Cited alongside, same era.
Tree-based batch mode reinforcement learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Cited alongside, same era.
Finite time bounds for sampling based fitted value iteration
Csaba Szepesvári and Rémi Munos · 2005
Cited alongside, same era.
Bandit based monte-carlo planning
Levente Kocsis and Csaba Szepesvári · 2006
Cited alongside, same era.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Cited alongside, same era.
The epoch-greedy algorithm for multi-armed bandits with side information
John Langford and Tong Zhang · 2008
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Later among the works it cites.
Deep learning for real-time atari game play using offline monte-carlo tree search planning
Xiaoxiao Guo, Satinder Singh, Honglak Lee, Richard L Lewis, and Xiaoshi Wang · 2014
Later among the works it cites.
Reinforcement and imitation learning via interactive no-regret learning
Stephane Ross and J Andrew Bagnell · 2014
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Later among the works it cites.
Contextual decision processes with low Bellman rank are PAC-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E. Schapire · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Error Propagation for Approximate Policy and Value Iteration
Amir-massoud Farahmand, Csaba Szepesvári, and Rémi Munos · 2010
Cited alongside, same era.
Monte-Carlo planning in large POMDPs
David Silver and Joel Veness · 2010
Cited alongside, same era.
Algorithms for reinforcement learning
Csaba Szepesvári · 2010
Cited alongside, same era.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Cited alongside, same era.
Sbeed: Convergent reinforcement learning with nonlinear function approximation
Bo Dai, Albert Shaw, Lihong Li, Lin Xiao, Niao He, Zhen Liu, Jianshu Chen, and Le Song · 2018
Later among the works it cites.
On Oracle-Efficient PAC RL with Rich Observations
Christoph Dann, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Closest in time.
Diagnosing bottlenecks in deep q-learning algorithms
Justin Fu, Aviral Kumar, Matthew Soh, and Sergey Levine · 2019
Closest in time.