Fetching the paper…
Reading the bibliography…
In this work, we present a new model-free and off-policy reinforcement learning (RL) algorithm, that is capable of finding a near-optimal policy with state-action observations from arbitrary behavior policies.
On tail probabilities for martingales
David A. Freedman · 1975
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S. Sutton · 1988
Earlier work this paper cites.
Q-learning
Christopher J. C. H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
On-line Q-learning using connectionist systems , volume 37
Gavin A. Rummery and Mahesan Niranjan · 1994
Earlier work this paper cites.
Residual algorithms: reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
Neuro-dynamic programming
Dimitri P. Bertsekas and John N. Tsitsiklis · 1996
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y. Ng and Stuart J. Russell · 2000
Earlier work this paper cites.
Valuing american options by simulation: A simple least-squares approach
Francis A. Longstaff and Eduardo S. Schwartz · 2001
Earlier work this paper cites.
Off-policy temporal-difference learning with function approximation
Doina Precup, Richard S. Sutton, and Sanjoy Dasgupta · 2001
Earlier work this paper cites.
Pricing in agent economies using multi-agent Q-learning
Gerald Tesauro and Jeffrey O. Kephart · 2002
Earlier work this paper cites.
Convex analysis and optimization
Dimitri P. Bertsekas, Angelia Nedić, and Asuman E. Ozdaglar · 2003
Cited alongside, same era.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Cited alongside, same era.
Matrix algebra
James E. Gentle · 2007
Cited alongside, same era.
Subgradient methods for saddle-point problems
Angelia Nedić and Asuman Ozdaglar · 2009
Cited alongside, same era.
Reinforcement learning in finite MDPs: PAC analysis
Alexander L. Strehl, Lihong Li, and Michael L. Littman · 2009
Cited alongside, same era.
Hoeffding’s inequality for supermartingales
Xiequan Fan, Ion Grama, and Quansheng Liu · 2012
Cited alongside, same era.
Concentration inequalities for sums and martingales
Bernard Bercu, Bernard Delyon, and Emmanuel Rio · 2015
Later among the works it cites.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández · 2015
Later among the works it cites.
Distributed policy evaluation under multiple behavior strategies
Sergio Valcarcel Macua, Jianshu Chen, Santiago Zazo, and Ali H. Sayed · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, et al · 2015
Later among the works it cites.
Stochastic primal-dual methods and sample complexity of reinforcement learning
Yichen Chen and Mengdi Wang · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sridhar Mahadevan and Bo Liu · 2012
Cited alongside, same era.
QD-learning: a collaborative distributed strategy for multi-agent reinforcement learning through consensus + + innovations
Soummya Kar, José MF Moura, and H. Vincent Poor · 2013
Cited alongside, same era.
Adventures in stochastic processes
Sidney I. Resnick · 2013
Cited alongside, same era.
Proximal reinforcement learning: A new theory of sequential decision making in primal-dual spaces
Sridhar Mahadevan, Bo Liu, Philip Thomas, Will Dabney, Steve Giguere, Nicholas Jacek, Ian Gemp, and Ji Liu · 2014
Cited alongside, same era.
Markov decision processes: Discrete stochastic dynamic programming
Martin L. Puterman · 2014
Cited alongside, same era.
Boosting the actor with dual critic
Bo Dai, Albert Shaw, Niao He, Lihong Li, and Le Song
Cited in the paper.
Mengdi Wang and Yichen Chen · 2016
Later among the works it cites.
Socially aware motion planning with deep reinforcement learning
Yu Fan Chen, Michael Everett, Miao Liu, and Jonathan P. How · 2017
Later among the works it cites.
Learning from conditional distributions via dual embeddings
Bo Dai, Niao He, Yunpeng Pan, Byron Boots, and Le Song · 2017
Later among the works it cites.
Primal-dual algorithm for distributed reinforcement learning: distributed GTD
Donghwan Lee, Hyungjin Yoon, and Naira Hovakimyan · 2018
Closest in time.
Fully decentralized multi-agent reinforcement learning with networked agents
Kaiqing Zhang, Zhuoran Yang, Han Liu, Tong Zhang, and Tamer Başar · 2018
Closest in time.
Stochastic primal-dual Q-learning algorithm for discounted MDPs
Donghwan Lee and Niao He · 2019
Closest in time.