Fetching the paper…
Reading the bibliography…
Off-policy evaluation (OPE) holds the promise of being able to leverage large, offline datasets for both evaluating and selecting complex policies for decision making.
Temporal difference learning and td-gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Planning treatment of ischemic heart disease with partially observable markov decision processes
Milos Hauskrecht and Hamish Fraser · 2000
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup · 2000
Earlier work this paper cites.
Scaling reinforcement learning toward robocup soccer
Peter Stone and Richard S Sutton · 2001
Earlier work this paper cites.
Reinforcement learning benchmarks and bake-offs ii
Alain Dutech, Timothy Edmunds, Jelle Kok, Michail Lagoudakis, Michael Littman, Martin Riedmiller, Bryan Russell, Bruno Scherrer, Richard Sutton, Stephan Timmer, et al · 2005
Earlier work this paper cites.
Agent based decision support system using reinforcement learning under emergency circumstances
Devinder Thapa, In-Sung Jung, and Gi-Nam Wang · 2005
Earlier work this paper cites.
Planning with approximate and learned models of markov decision processes
Cosmin Paduraru · 2007
Earlier work this paper cites.
Evaluation of policy gradient methods and variants on the cart-pole benchmark
Martin Riedmiller, Jan Peters, and Stefan Schaal · 2007
Earlier work this paper cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Richard S Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvári, and Eric Wiewiora · 2009
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E Schapire · 2010
Earlier work this paper cites.
Fitted policy search: Direct policy search using a batch reinforcement learning approach
Martino Migliavacca, Alessio Pecorino, Matteo Pirotta, Marcello Restelli, and Andrea Bonarini · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Counterfactual reasoning and learning systems: The example of computational advertising
Léon Bottou, Jonas Peters, Joaquin Quiñonero-Candela, Denis X Charles, D Max Chickering, Elon Portugaly, Dipankar Ray, Patrice Simard, and Ed Snelson · 2013
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Doubly robust policy evaluation and optimization
Miroslav Dudík, Dumitru Erhan, John Langford, Lihong Li, et al · 2014
Earlier work this paper cites.
Offline policy evaluation across representations with applications to educational games
Travis Mandel, Yun-En Liu, Sergey Levine, Emma Brunskill, and Zoran Popovic · 2014
Earlier work this paper cites.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Counterfactual risk minimization: Learning from logged bandit feedback
Adith Swaminathan and Thorsten Joachims · 2015
Cited alongside, same era.
Personalized ad recommendation systems for life-time value optimization with guarantees
Georgios Theocharous, Philip S Thomas, and Mohammad Ghavamzadeh · 2015
Cited alongside, same era.
High-confidence off-policy evaluation
Philip S Thomas, Georgios Theocharous, and Mohammad Ghavamzadeh · 2015
Cited alongside, same era.
Recommender systems , volume 1
Charu C Aggarwal et al · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Behaviour policy estimation in off-policy policy evaluation: Calibration matters
Aniruddh Raghu, Omer Gottesman, Yao Liu, Matthieu Komorowski, Aldo Faisal, Finale Doshi-Velez, and Emma Brunskill · 2018
Later among the works it cites.
Off-policy evaluation and learning from logged bandit feedback: Error reduction via surrogate policy
Yuan Xie, Boyi Liu, Qiang Liu, Zhaoran Wang, Yuan Zhou, and Jian Peng · 2018
Later among the works it cites.
Importance sampling policy evaluation with an estimated behavior policy
Josiah Hanna, Scott Niekum, and Peter Stone · 2019
Later among the works it cites.
Off-policy evaluation via off-policy classification
Alexander Irpan, Kanishka Rao, Konstantinos Bousmalis, Chris Harris, Julian Ibarz, and Sergey Levine · 2019
Later among the works it cites.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc G. Bellemare · 2016
Cited alongside, same era.
An emphatic approach to the problem of off-policy temporal-difference learning
Richard S Sutton, A Rupam Mahmood, and Martha White · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip Thomas and Emma Brunskill · 2016
Cited alongside, same era.
Importance sampling for fair policy selection
Shayan Doroudi, Philip S Thomas, and Emma Brunskill · 2017
Cited alongside, same era.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
Shixiang Gu, Ethan Holly, Timothy Lillicrap, and Sergey Levine · 2017
Cited alongside, same era.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
Aravind Rajeswaran, Vikash Kumar, Abhishek Gupta, Giulia Vezzani, John Schulman, Emanuel Todorov, and Sergey Levine · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Learning to drive in a day
Alex Kendall, Jeffrey Hawke, David Janz, Przemyslaw Mazur, Daniele Reda, John-Mark Allen, Vinh-Dieu Lam, Alex Bewley, and Amar Shah · 2019
Later among the works it cites.
Batch policy learning under constraints
Hoang M Le, Cameron Voloshin, and Yisong Yue · 2019
Later among the works it cites.
Learning when-to-treat policies
Xinkun Nie, Emma Brunskill, and Stefan Wager · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Later among the works it cites.
Empirical study of off-policy policy evaluation for reinforcement learning
Cameron Voloshin, Hoang M Le, Nan Jiang, and Yisong Yue · 2019
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Rl unplugged: Benchmarks for offline reinforcement learning
Caglar Gulcehre, Ziyu Wang, Alexander Novikov, Tom Le Paine, Sergio Gómez Colmenarejo, Konrad Zolna, Rishabh Agarwal, Josh Merel, Daniel Mankowitz, Cosmin Paduraru, et al · 2020
Later among the works it cites.
Acme: A research framework for distributed reinforcement learning
Matt Hoffman, Bobak Shahriari, John Aslanides, Gabriel Barth-Maron, Feryal Behbahani, Tamara Norman, Abbas Abdolmaleki, Albin Cassirer, Fan Yang, Kate Baumli, et al · 2020
Later among the works it cites.
Statistical bootstrapping for uncertainty estimation in off-policy evaluation, 2020
Ilya Kostrikov and Ofir Nachum · 2020
Later among the works it cites.
Hyperparameter selection for offline reinforcement learning
Tom Le Paine, Cosmin Paduraru, Andrea Michi, Caglar Gulcehre, Konrad Zolna, Alexander Novikov, Ziyu Wang, and Nando de Freitas · 2020
Later among the works it cites.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
Noah Y Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki, Michael Neunert, Thomas Lampe, Roland Hafner, and Martin Riedmiller · 2020
Later among the works it cites.
Ziyu Wang, Alexander Novikov, Konrad Żołna, Jost Tobias Springenberg, Scott Reed, Bobak Shahriari, Noah Siegel, Josh Merel, Caglar Gulcehre, Nicolas Heess, and Nando de Freitas · 2020
Later among the works it cites.
Batch stationary distribution estimation
Junfeng Wen, Bo Dai, Lihong Li, and Dale Schuurmans · 2020
Later among the works it cites.
Off-policy evaluation via the regularized lagrangian
Mengjiao Yang, Ofir Nachum, Bo Dai, Lihong Li, and Dale Schuurmans · 2020
Later among the works it cites.
Autoregressive dynamics models for offline policy evaluation and optimization
Michael R Zhang, Thomas Paine, Ofir Nachum, Cosmin Paduraru, George Tucker, ziyu wang, and Mohammad Norouzi · 2021
Closest in time.