Fetching the paper…
Reading the bibliography…
We show that the popular reinforcement learning (RL) strategy of estimating the state-action value (Q-function) by minimizing the mean squared Bellman error leads to a regression problem with confounding, the inputs and output noise being correlated.
Empirical study of off-policy policy evaluation for reinforcement learning
Cameron Voloshin, Hoang M Le, Nan Jiang, and Yisong Yue · 1911
Earlier work this paper cites.
The Tariff on Animal and Vegetable Oils
P.G. Wright · 1928
Earlier work this paper cites.
Large sample properties of generalized method of moments estimators
Lars Peter Hansen · 1982
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
Andrew G Barto, Richard S Sutton, and Charles W Anderson · 1983
Earlier work this paper cites.
Lifetime earnings and the Vietnam era draft lottery: Evidence from social security administrative records
Joshua D. Angrist · 1990
Earlier work this paper cites.
Efficient Memory-Based Learning for Robot Control
Andrew William Moore · 1990
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton · 1990
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
Identification of causal effects using instrumental variables
Joshua D. Angrist, Guido W. Imbens, and Donald B. Rubin · 1996
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Steven J Bradtke and Andrew G Barto · 1996
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup · 2000
Earlier work this paper cites.
Off-policy temporal-difference learning with function approximation
Doina Precup, Richard S Sutton, and Sanjoy Dasgupta · 2001
Earlier work this paper cites.
Least-squares policy iteration
Michail G Lagoudakis and Ronald Parr · 2003
Earlier work this paper cites.
Instrumental variable estimation of nonparametric models
Whitney K. Newey and James L. Powell · 2003
Earlier work this paper cites.
Retrospectives: Who invented instrumental variable regression?
James H. Stock and Francesco Trebbi · 2003
Earlier work this paper cites.
Reinforcement learning with gaussian processes
Yaakov Engel, Shie Mannor, and Ron Meir · 2005
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Earlier work this paper cites.
Linear inverse problems in structural econometrics estimation based on spectral decomposition and regularization
Marine Carrasco, Jean-Pierre Florens, and Eric Renault · 2007
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Real-time reinforcement learning by sequential actor–critics and experience replay
Paweł Wawrzyński · 2009
Cited alongside, same era.
Nonparametric instrumental regression
S. Darolles, Y. Fan, J. P. Florens, and E. Renault · 2011
Cited alongside, same era.
Doubly robust policy evaluation and learning
Miroslav Dudík, John Langford, and Lihong Li · 2011
Cited alongside, same era.
Measuring the price responsiveness of gasoline demand: Economic shape restrictions and nonparametric demand estimation
Richard Blundell, Joel Horowitz, and Matthias Parey · 2012
Cited alongside, same era.
Synthesis and stabilization of complex behaviors through online trajectory optimization
Yuval Tassa, Tom Erez, and Emanuel Todorov · 2012
Cited alongside, same era.
Behaviour suite for reinforcement learning
Ian Osband, Yotam Doron, Matteo Hessel, John Aslanides, Eren Sezener, Andre Saraiva, Katrina McKinney, Tor Lattimore, Csaba Szepesvari, Satinder Singh, et al · 2019
Later among the works it cites.
Deterministic Bellman residual minimization
Ehsan Saleh and Nan Jiang · 2019
Later among the works it cites.
Autonomous navigation of stratospheric balloons using reinforcement learning
Marc G Bellemare, Salvatore Candido, Pablo Samuel Castro, Jun Gong, Marlos C Machado, Subhodeep Moitra, Sameera S Ponda, and Ziyu Wang · 2020
Later among the works it cites.
Minimax estimation of conditional moment models
Nishanth Dikkala, Greg Lewis, Lester Mackey, and Vasilis Syrgkanis · 2020
Later among the works it cites.
Rl unplugged: A suite of benchmarks for offline reinforcement learning
Caglar Gulcehre, Ziyu Wang, Alexander Novikov, Tom Le Paine, Sergio Gomez Colmenarejo, Konrad Zolna, Rishabh Agarwal, Josh Merel, Daniel Mankowitz, Cosmin Paduraru, et al · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip Thomas and Emma Brunskill · 2016
Cited alongside, same era.
Distributed distributional deterministic policy gradients
Gabriel Barth-Maron, Matthew W Hoffman, David Budden, Will Dabney, Dan Horgan, TB Dhruva, Alistair Muldal, Nicolas Heess, and Timothy Lillicrap · 2018
Cited alongside, same era.
Later among the works it cites.
Acme: A research framework for distributed reinforcement learning
Matt Hoffman, Bobak Shahriari, John Aslanides, Gabriel Barth-Maron, Feryal Behbahani, Tamara Norman, Abbas Abdolmaleki, Albin Cassirer, Fan Yang, Kate Baumli, Sarah Henderson, Alex Novikov, Sergio Gómez Colmenarejo, Serkan Cabi, Caglar Gulcehre, Tom Le Paine, Andrew Cowie, Ziyu Wang, Bilal Piot, and Nando de Freitas · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Later among the works it cites.
Provably efficient neural estimation of structural equation model: An adversarial approach
Liao Luofeng, Chen You-Lin, Yang Zhuoran, Dai Bo, Wang Zhaoran, and Kolar Mladen · 2020
Later among the works it cites.
Black-box off-policy estimation for infinite-horizon reinforcement learning
Ali Mousavi, Lihong Li, Qiang Liu, and Denny Zhou · 2020
Later among the works it cites.
Hyperparameter selection for offline reinforcement learning
Tom Le Paine, Cosmin Paduraru, Andrea Michi, Caglar Gulcehre, Konrad Zolna, Alexander Novikov, Ziyu Wang, and Nando de Freitas · 2020
Later among the works it cites.
dm_control: Software and tasks for continuous control, 2020
Yuval Tassa, Saran Tunyasuvunakool, Alistair Muldal, Yotam Doron, Siqi Liu, Steven Bohez, Josh Merel, Tom Erez, Timothy Lillicrap, and Nicolas Heess · 2020
Later among the works it cites.
Minimax weight and Q-function learning for off-policy evaluation
Masatoshi Uehara, Jiawei Huang, and Nan Jiang · 2020
Later among the works it cites.
Off-policy evaluation via the regularized Lagrangian
Mengjiao Yang, Ofir Nachum, Bo Dai, Lihong Li, and Dale Schuurmans · 2020
Later among the works it cites.
Off-policy evaluation in infinite-horizon reinforcement learning with latent confounders
Andrew Bennett, Nathan Kallus, Lihong Li, and Ali Mousavi · 2021
Closest in time.
Benchmarks for deep off-policy evaluation
Justin Fu, Mohammad Norouzi, Ofir Nachum, George Tucker, ziyu wang, Alexander Novikov, Mengjiao Yang, Michael R Zhang, Yutian Chen, Aviral Kumar, Cosmin Paduraru, Sergey Levine, and Thomas Paine · 2021
Closest in time.
Causal reinforcement learning: An instrumental variable approach
Jin Li, Ye Luo, and Xiaowei Zhang · 2021
Closest in time.
Instrumental variable value iteration for causal offline reinforcement learning
Luofeng Liao, Zuyue Fu, Zhuoran Yang, Mladen Kolar, and Zhaoran Wang · 2021
Closest in time.
Learning deep features in instrumental variable regression
Liyuan Xu, Yutian Chen, Siddarth Srinivasan, Nando de Freitas, Arnaud Doucet, and Arthur Gretton · 2021
Closest in time.
Autoregressive dynamics models for offline policy evaluation and optimization
Michael R Zhang, Thomas Paine, Ofir Nachum, Cosmin Paduraru, George Tucker, ziyu wang, and Mohammad Norouzi · 2021
Closest in time.