Fetching the paper…
Reading the bibliography…
A central object of study in Reinforcement Learning (RL) is the Markovian policy, in which an agent's actions are chosen from a memoryless probability distribution, conditioned only on its current state.
An Introduction to Probability Theory and Its Applications
W. Feller · 1968
Earlier work this paper cites.
Reconstructing human skill with machine learning
T. Urbancic · 1994
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
Importance sampling for reinforcement learning with multiple objectives
C. R. Shelton · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
S. Kakade and J. Langford · 2002
Earlier work this paper cites.
Learning options in reinforcement learning
M. Stolle and D. Precup · 2002
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
A. G. Barto and S. Mahadevan · 2003
Earlier work this paper cites.
On non-stationary policies and maximal invariant safe sets of controlled markov chains
W. Wu, A. Arapostathis, and R. Kumar · 2004
Earlier work this paper cites.
Ensemble algorithms in reinforcement learning
M. A. Wiering and H. Van Hasselt · 2008
Earlier work this paper cites.
Constructing stochastic mixture policies for episodic multiobjective reinforcement learning tasks
P. Vamplew, R. Dazeley, E. Barker, and A. Kelarev · 2009
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Batch Reinforcement Learning
S. Lange, T. Gabel, and M. Riedmiller · 2012
Earlier work this paper cites.
On the use of non-stationary policies for stationary infinite-horizon markov decision processes
B. Scherrer and B. Lesner · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. A. Riedmiller · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
High confidence policy improvement
P. S. Thomas, G. Theocharous, and M. Ghavamzadeh · 2015
Earlier work this paper cites.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Imitation learning: A survey of learning methods
A. Hussein, M. M. Gaber, E. Elyan, and C. Jayne · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Cited alongside, same era.
Mix & match agent curricula for reinforcement learning
W. Czarnecki, S. Jayakumar, M. Jaderberg, L. Hasenclever, Y. W. Teh, N. Heess, S. Osindero, and R. Pascanu · 2018
Cited alongside, same era.
Distributed prioritized experience replay
D. Horgan, J. Quan, D. Budden, G. Barth-Maron, M. Hessel, H. van Hasselt, and D. Silver · 2018
Cited alongside, same era.
Reinforcement learning algorithm selection
R. Laroche and R. Féraud · 2018
An optimistic perspective on offline reinforcement learning
R. Agarwal, D. Schuurmans, and M. Norouzi · 2020
Later among the works it cites.
The importance of pessimism in fixed-dataset policy optimization
J. Buckman, C. Gelada, and M. G. Bellemare · 2020
Later among the works it cites.
A reduction from reinforcement learning to no-regret online learning
C.-A. Cheng, R. T. des Combes, B. Boots, and G. Gordon · 2020
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Data-efficient hierarchical reinforcement learning
O. Nachum, S. S. Gu, H. Lee, and S. Levine · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Cited alongside, same era.
Behavioral cloning from observation
F. Torabi, G. Warnell, and P. Stone · 2018
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine · 2019
Cited alongside, same era.
Benchmarking batch deep reinforcement learning algorithms
S. Fujimoto, E. Conti, M. Ghavamzadeh, and J. Pineau · 2019
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2019
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems, 2020
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Later among the works it cites.
Off-policy actor-critic with shared experience replay
S. Schmitt, M. Hessel, and K. Simonyan · 2020
Later among the works it cites.
Safe policy improvement with an estimated baseline policy
T. D. Simao, R. Laroche, and R. Tachet des Combes · 2020
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch · 2021
Later among the works it cites.
Dr Jekyll and Mr Hyde: The strange case of off-policy policy updates
R. Laroche and R. Tachet des Combes · 2021
Later among the works it cites.
Multi-objective spibb: Seldonian offline policy improvement with safety constraints in finite mdps
H. Satija, P. S. Thomas, J. Pineau, and R. Laroche · 2021
Later among the works it cites.
Near-optimal offline reinforcement learning via double variance reduction
M. Yin, Y. Bai, and Y.-X. Wang · 2021
Later among the works it cites.
COMBO: Conservative offline model-based policy optimization
T. Yu, A. Kumar, R. Rafailov, A. Rajeswaran, S. Levine, and C. Finn · 2021
Later among the works it cites.
Rvs: What is essential for offline RL via supervised learning?
S. Emmons, B. Eysenbach, I. Kostrikov, and S. Levine · 2022
Closest in time.
Generalized decision transformer for offline hindsight information matching
H. Furuta, Y. Matsuo, and S. S. Gu · 2022
Closest in time.
Safe policy improvement approaches on discrete markov decision processes, 2022
P. Scholl, F. Dietrich, C. Otte, and S. Udluft · 2022
Closest in time.
Pessimistic q-learning for offline reinforcement learning: Towards optimal sample complexity
L. Shi, G. Li, Y. Wei, Y. Chen, and Y. Chi · 2022
Closest in time.
Theoretical foundations of reinforcement learning course, 2022
C. Szepesvári · 2022
Closest in time.