Fetching the paper…
Reading the bibliography…
In many real-world reinforcement learning applications, access to the environment is limited to a fixed dataset, instead of direct (online) interaction with the environment.
Monte carlo sampling methods using markov chains and their applications
W Keith Hastings · 1970
Earlier work this paper cites.
Convergence of Stochastic Processes
D Pollard · 1984
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Martin L Puterman · 1994
Earlier work this paper cites.
Rates of convergence for empirical processes of stationary mixing sequences
Bin Yu · 1994
Earlier work this paper cites.
Sphere packing numbers for subsets of the boolean n-cube with bounded vapnik-chervonenkis dimension
David Haussler · 1995
Earlier work this paper cites.
Inferences for case-control and semiparametric two-sample density ratio models
Jing Qin · 1998
Earlier work this paper cites.
Introduction to reinforcement learning
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Hierarchical reinforcement learning with the MAXQ value function decomposition
Thomas G Dietterich · 2000
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
D. Precup, R. Sutton, and S. Singh · 2000
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup · 2000
Earlier work this paper cites.
Marginal mean models for dynamic regimes
Susan A Murphy, Mark J van der Laan, James M Robins, and Conduct Problems Prevention Research Group · 2001
Earlier work this paper cites.
Off-policy temporal difference learning with function approximation
D. Precup, R. Sutton, and S. Dasgupta · 2001
Earlier work this paper cites.
Dynamic Programming
Richard Ernest Bellman · 2003
Earlier work this paper cites.
Semiparametric density estimation under a two-sample density ratio model
Kuang Fu Cheng, Chih-Kang Chu, et al · 2004
Earlier work this paper cites.
Discriminative learning for differing training and test distributions
Steffen Bickel, Michael Brückner, and Tobias Scheffer · 2007
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Direct importance estimation for covariate shift adaptation
Masashi Sugiyama, Taiji Suzuki, Shinichi Nakajima, Hisashi Kashima, Paul von Bünau, and Motoaki Kawanabe · 2008
Earlier work this paper cites.
Dyna-style planning with linear function approximation and prioritized sweeping
Richard S Sutton, Csaba Szepesvári, Alborz Geramifard, and Michael Bowling · 2008
Earlier work this paper cites.
Covariate shift by kernel mean matching
Arthur Gretton, Alex J Smola, Jiayuan Huang, Marcel Schmittfull, Karsten M Borgwardt, and Bernhard Schöllkopf · 2009
Cited alongside, same era.
A least-squares approach to direct importance estimation
Takafumi Kanamori, Shohei Hido, and Masashi Sugiyama · 2009
Cited alongside, same era.
Variational analysis
R Tyrrell Rockafellar and Roger J-B Wets · 2009
Cited alongside, same era.
Lectures on stochastic programming: modeling and theory
Alexander Shapiro, Darinka Dentcheva, and Andrzej Ruszczyński · 2009
Cited alongside, same era.
Estimating divergence functionals and the likelihood ratio by convex risk minimization
XuanLong Nguyen, Martin J Wainwright, and Michael I Jordan · 2010
Cited alongside, same era.
Doubly robust policy evaluation and learning
Miroslav Dudík, John Langford, and Lihong Li · 2011
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Later among the works it cites.
Learning from conditional distributions via dual embeddings
Bo Dai, Niao He, Yunpeng Pan, Byron Boots, and Le Song · 2016
Later among the works it cites.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Later among the works it cites.
Deep reinforcement learning for dialogue generation
Jiwei Li, Will Monroe, Alan Ritter, Michel Galley, Jianfeng Gao, and Dan Jurafsky · 2016
Later among the works it cites.
An emphatic approach to the problem of off-policy temporal-difference learning
Richard S Sutton, A Rupam Mahmood, and Martha White · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms
Lihong Li, Wei Chu, John Langford, and Xuanhui Wang · 2011
Cited alongside, same era.
Finite-sample analysis of least-squares policy iteration
Alessandro Lazaric, Mohammad Ghavamzadeh, and Rémi Munos · 2012
Cited alongside, same era.
Density ratio estimation in machine learning
Masashi Sugiyama, Taiji Suzuki, and Takafumi Kanamori · 2012
Cited alongside, same era.
Batch mode reinforcement learning based on the synthesis of artificial trajectories
Raphael Fonteneau, Susan A. Murphy, Louis Wehenkel, and Damien Ernst · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Off-policy Evaluation in Markov Decision Processes
C. Paduraru · 2013
Cited alongside, same era.
Later among the works it cites.
Data-efficient off-policy policy evaluation for reinforcement learning
P. Thomas and E. Brunskill · 2016
Later among the works it cites.
Sbeed: Convergent reinforcement learning with nonlinear function approximation
Bo Dai, Albert Shaw, Lihong Li, Lin Xiao, Niao He, Zhen Liu, Jianshu Chen, and Le Song · 2017
Later among the works it cites.
Stochastic variance reduction methods for policy evaluation
Simon S Du, Jianshu Chen, Lihong Li, Lin Xiao, and Dengyong Zhou · 2017
Later among the works it cites.
Consistent on-line off-policy evaluation
Assaf Hallak and Shie Mannor · 2017
Later among the works it cites.
Off-policy evaluation for slate recommendation
A. Swaminathan, A. Krishnamurthy, A. Agarwal, M. Dudík, J. Langford, D. Jose, and I. Zitouni · 2017
Later among the works it cites.
Optimal and adaptive off-policy evaluation in contextual bandits
Yu-Xiang Wang, Alekh Agarwal, and Miroslav Dudik · 2017
Later among the works it cites.
Learning dexterous in-hand manipulation
Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafal Jozefowicz, Bob McGrew, Jakub Pachocki, Arthur Petron, Matthias Plappert, Glenn Powell, Alex Ray, et al · 2018
Later among the works it cites.
More robust doubly robust off-policy evaluation
Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh · 2018
Later among the works it cites.
Off-policy deep reinforcement learning by bootstrapping the covariate shift
Carles Gelada and Marc G Bellemare · 2018
Later among the works it cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Qiang Liu, Lihong Li, Ziyang Tang, and Dengyong Zhou · 2018
Later among the works it cites.
Neural approaches to Conversational AI
Jianfeng Gao, Michel Galley, and Lihong Li · 2019
Closest in time.
Off-policy policy gradient with state distribution correction
Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill · 2019
Closest in time.