Fetching the paper…
Reading the bibliography…
In batch reinforcement learning (RL), one often constrains a learned policy to be close to the behavior (data-generating) policy, e.g., by constraining the learned action distribution to differ from the behavior policy by some maximum degree that is the same at each state.
Infinite-dimensional quadratic optimization: interior-point methods and control applications
L. Faybusovich and J. Moore · 1997
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
S. Kakade and J. Langford · 2002
Earlier work this paper cites.
The linear programming approach to approximate dynamic programming
P. De Farias and B. Van Roy · 2003
Earlier work this paper cites.
Convex optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
A tutorial on MM algorithms
D. Hunter and K. Lange · 2004
Earlier work this paper cites.
Weighted Csiszár-Kullback-Pinsker inequalities and applications to transportation inequalities
F. Bolley and C. Villani · 2005
Earlier work this paper cites.
A kernel approach to comparing distributions
A. Gretton, K. Borgwardt, M. Rasch, B. Schölkopf, and A. Smola · 2007
Earlier work this paper cites.
Lectures on stochastic programming: modeling and theory
A. Shapiro, D. Dentcheva, and A. Ruszczyński · 2009
Earlier work this paper cites.
Imitation and reinforcement learning
J. Kober and J. Peters · 2010
Earlier work this paper cites.
Batch reinforcement learning
S. Lange, T. Gabel, and M. Riedmiller · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
Safe policy iteration
M. Pirotta, M. Restelli, A. Pecorino, and D. Calandriello · 2013
Earlier work this paper cites.
A block coordinate descent method for regularized multiconvex optimization with applications to nonnegative tensor factorization and completion
Y. Xu and W. Yin · 2013
Cited alongside, same era.
Weighted importance sampling for off-policy learning with linear function approximation
A. Mahmood, H. van Hasselt, and R. Sutton · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
T. Lillicrap, J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Cited alongside, same era.
A. Rusu, S. Colmenarejo, C. Gulcehre, G. Desjardins, J. Kirkpatrick, R. Pascanu, V. Mnih, K. Kavukcuoglu, and R. Hadsell · 2015
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Cited alongside, same era.
Horizon: Facebook’s open source applied reinforcement learning platform
J. Gauci, E. Conti, Y. Liang, K. Virochsiri, Y. He, Z. Kaden, V. Narayanan, and X. Ye · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
Natural gradient deep Q-learning
E. Knight and O. Lerner · 2018
Later among the works it cites.
T. Silver, K. Allen, J. Tenenbaum, and L. Kaelbling · 2018
Later among the works it cites.
Reinforcement learning: An introduction
R. Sutton and A. Barto · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
OpenAI Gym, 2016
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
Safe policy improvement by minimizing robust baseline regret
M. Ghavamzadeh, M. Petrik, and Y. Chow · 2016
Cited alongside, same era.
Deep reinforcement learning with double Q-learning
H. Van Hasselt, A. Guez, and D. Silver · 2016
Cited alongside, same era.
Combining policy gradient and Q-learning
B. O’Donoghue, R. Munos, K. Kavukcuoglu, and V. Mnih · 2016
Cited alongside, same era.
Optnet: Differentiable optimization as a layer in neural networks
B. Amos and Z. Kolter · 2017
Cited alongside, same era.
Safe policy improvement with baseline bootstrapping
R. Laroche and P. Trichelair · 2017
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2018
Cited alongside, same era.
Later among the works it cites.
Large-scale Markov decision problems via the linear programming dual
Y. Abbasi-Yadkori, P. Bartlett, X. Chen, and A. Malek · 2019
Later among the works it cites.
Benchmarking batch deep reinforcement learning algorithms
S. Fujimoto, E. Conti, M. Ghavamzadeh, and J. Pineau · 2019
Later among the works it cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
N. Jaques, A. Ghandeharioun, J. Shen, C. Ferguson, A. Lapedriza, N. Jones, S. Gu, and R. Picard · 2019
Later among the works it cites.
Residual reinforcement learning for robot control
T. Johannink, S. Bahl, A. Nair, J. Luo, A. Kumar, M. Loskyll, J. Ojea, E. Solowjow, and S.Levine · 2019
Later among the works it cites.
Stabilizing off-policy Q-learning via bootstrapping error reduction
A. Kumar, J. Fu, G. Tucker, and S. Levine · 2019
Later among the works it cites.
Multi-agent manipulation via locomotion using hierarchical sim2real
O. Nachum, M. Ahn, H. Ponte, S. Gu, and V. Kumar · 2019
Later among the works it cites.