Fetching the paper…
Reading the bibliography…
In many real-world applications of reinforcement learning (RL), interactions with the environment are limited due to cost or feasibility.
Multi-agent manipulation via locomotion using hierarchical sim2real
Ofir Nachum, Michael Ahn, Hugo Ponte, Shixiang Gu, and Vikash Kumar · 1908
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L Puterman · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
Nonlinear Programming
D. P. Bertsekas · 1999
Earlier work this paper cites.
Convex analysis and variational problems , volume 28
Ivar Ekeland and Roger Temam · 1999
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
On the existence of fixed points for approximate value iteration and temporal-difference learning
Daniela Pucci de Farias and Benjamin Van Roy · 2000
Earlier work this paper cites.
Actor-critic algorithms
Vijay R. Konda and John N. Tsitsiklis · 2000
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
D. Precup, R. Sutton, and S. Singh · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Linearly-solvable Markov decision problems
Emanuel Todorov · 2006
Earlier work this paper cites.
Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Dual representations for dynamic programming
Tao Wang, Daniel Lizotte, Michael Bowling, and Dale Schuurmans · 2008
Earlier work this paper cites.
Estimating divergence functionals and the likelihood ratio by convex risk minimization
XuanLong Nguyen, Martin J Wainwright, and Michael I Jordan · 2010
Earlier work this paper cites.
Thomas Degris, Martha White, and Richard S Sutton · 2012
Earlier work this paper cites.
Markov chains and stochastic stability
Sean P Meyn and Richard L Tweedie · 2012
Earlier work this paper cites.
Trading value and information in MDPs
Jonathan Rubin, Ohad Shamir, and Naftali Tishby · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Playing Atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Off-policy learning with eligibility traces: A survey
M. Geist and B. Scherrer · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Batch learning from logged bandit feedback through counterfactual risk minimization
Adith Swaminathan and Thorsten Joachims · 2015
Cited alongside, same era.
Randomized Linear Programming Solves the Discounted Markov Decision Problem In Nearly-Linear Running Time
Mengdi Wang · 2017
Later among the works it cites.
Learning dexterous in-hand manipulation
Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafal Jozefowicz, Bob McGrew, Jakub Pachocki, Arthur Petron, Matthias Plappert, Glenn Powell, Alex Ray, et al · 2018
Later among the works it cites.
Scalable bilinear π \pi learning using state and action features
Yichen Chen, Lihong Li, and Mengdi Wang · 2018
Later among the works it cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke van Hoof, and David Meger · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Stochastic primal-dual methods and sample complexity of reinforcement learning
Yichen Chen and Mengdi Wang · 2016
Cited alongside, same era.
Learning from conditional distributions via dual embeddings
Bo Dai, Niao He, Yunpeng Pan, Byron Boots, and Le Song · 2016
Cited alongside, same era.
Regularized policy iteration with nonparametric function spaces
Amir-massoud Farahmand, Mohammad Ghavamzadeh, Csaba Szepesvári, and Shie Mannor · 2016
Cited alongside, same era.
Taming the noise in reinforcement learning via soft updates
Roy Fox, Ari Pakman, and Naftali Tishby · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
R. Munos, T. Stepleton, A. Harutyunyan, and M. Bellemare · 2016
Cited alongside, same era.
Sample efficient actor-critic with experience replay
Ziyu Wang, Victor Bapst, Nicolas Heess, Volodymyr Mnih, Remi Munos, Koray Kavukcuoglu, and Nando de Freitas · 2016
Cited alongside, same era.
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado Van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2018
Later among the works it cites.
Stochastic primal-dual q-learning
Donghwan Lee and Niao He · 2018
Later among the works it cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Qiang Liu, Lihong Li, Ziyang Tang, and Dengyong Zhou · 2018
Later among the works it cites.
Non-delusional Q-learning and value-iteration
Tyler Lu, Dale Schuurmans, and Craig Boutilier · 2018
Later among the works it cites.
Smoothed action value functions for learning Gaussian policies
Ofir Nachum, Mohammad Norouzi, George Tucker, and Dale Schuurmans · 2018
Later among the works it cites.
A kernel loss for solving the Bellman equation
Yihao Feng, Lihong Li, and Qiang Liu · 2019
Closest in time.
Neural approaches to Conversational AI
Jianfeng Gao, Michel Galley, and Lihong Li · 2019
Closest in time.
A theory of regularized Markov decision processes
Matthieu Geist, Bruno Scherrer, and Olivier Pietquin · 2019
Closest in time.
Imitation learning via off-policy distribution matching
Ilya Kostrikov, Ofir Nachum, and Jonathan Tompson · 2019
Closest in time.
Off-policy policy gradient with state distribution correction
Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill · 2019
Closest in time.
Doubly robust bias reduction in infinite horizon off-policy estimation
Ziyang Tang, Yihao Feng, Lihong Li, Dengyong Zhou, and Qiang Liu · 2019
Closest in time.
Minimax weight and Q-function learning for off-policy evaluation
Masatoshi Uehara and Nan Jiang · 2019
Closest in time.
GenDICE: Generalized offline estimation of stationary values, 2020
Ruiyi Zhang, Bo Dai, Li Lihong, and Dale Schuurmans · 2020
Closest in time.