Fetching the paper…
Reading the bibliography…
Off-policy evaluation (OPE) is the problem of evaluating new policies using historical data obtained from a different policy.
Inverse reinforcement learning for decentralized non-cooperative multiagent systems. In SMC . 1930–1935
Tummalapalli Sudhamsh Reddy, Vamsikrishna Gopikrishna, Gergely Zaruba, and Manfred Huber. 2012 · 1935
Earlier work this paper cites.
Non-cooperative games
John Nash. 1951 · 1951
Earlier work this paper cites.
Stochastic games
Lloyd S Shapley. 1953 · 1953
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman. 1994 · 1994
Earlier work this paper cites.
Estimation of regression coefficients when some regressors are not always observed
James M Robins, Andrea Rotnitzky, and Lue Ping Zhao. 1994 · 1994
Earlier work this paper cites.
A generalized reinforcement-learning model: Convergence and applications. In ICML . 310–318
Michael L Littman and Csaba Szepesvári. 1996 · 1996
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto. 1998 · 1998
Earlier work this paper cites.
Rademacher and Gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson. 2002 · 2002
Earlier work this paper cites.
Efficient estimation of average treatment effects using the estimated propensity score
Keisuke Hirano, Guido W Imbens, and Geert Ridder. 2003 · 2003
Earlier work this paper cites.
Nash Q-learning for general-sum stochastic games
Junling Hu and Michael P Wellman. 2003 · 2003
Earlier work this paper cites.
Optimal dynamic treatment regimes
Susan A Murphy. 2003 · 2003
Earlier work this paper cites.
Safe strategies for agent modelling in games. In AAAI Fall Symposium on Artificial Multi-agent Learning . 103–110
Peter McCracken and Michael Bowling. 2004 · 2004
Earlier work this paper cites.
Concentration inequalities and asymptotic results for ratio type empirical processes
Evarist Giné, Vladimir Koltchinskii, et al · 2006
Earlier work this paper cites.
Optimal unbiased estimators for evaluating agent performance. In AAAI . 573–579
Martin Zinkevich, Michael Bowling, Nolan Bard, Morgan Kan, and Darse Billings. 2006 · 2006
Earlier work this paper cites.
Semiparametric theory and missing data
Anastasios Tsiatis. 2007 · 2007
Earlier work this paper cites.
Strategy evaluation in extensive games with importance sampling. In ICML . 72–79
Michael Bowling, Michael Johanson, Neil Burch, and Duane Szafron. 2008 · 2008
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
Lucian Busoniu, Robert Babuska, and Bart De Schutter. 2008 · 2008
Earlier work this paper cites.
Regret minimization in games with incomplete information. In NeurIPS . 1729–1736
Martin Zinkevich, Michael Johanson, Michael Bowling, and Carmelo Piccione. 2008 · 2008
Earlier work this paper cites.
Data biased robust counter strategies. In AISTATS . 264–271
Michael Johanson and Michael Bowling. 2009 · 2009
Cited alongside, same era.
Effective short-term opponent exploitation in simplified poker
Finnegan Southey, Bret Hoehn, and Robert C Holte. 2009 · 2009
Cited alongside, same era.
Multi-agent inverse reinforcement learning. In ICMLA . 395–400
Sriraam Natarajan, Gautam Kunapuli, Kshitij Judah, Prasad Tadepalli, Kristian Kersting, and Jude Shavlik. 2010 · 2010
Cited alongside, same era.
Generalized Sampling and Variance in Counterfactual Regret Minimization.. In AAAI . 1355–1361
Richard G Gibson, Marc Lanctot, Neil Burch, Duane Szafron, and Michael Bowling. 2012 · 2012
Cited alongside, same era.
Online implicit agent modelling. In AAMAS . 255–262
Nolan Bard, Michael Johanson, Neil Burch, and Michael Bowling. 2013 · 2013
Cited alongside, same era.
Baseline: practical control variates for agent evaluation in zero-sum domains.. In AAMAS . 1005–1012
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Later among the works it cites.
Double/debiased machine learning for treatment and structural parameters
Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins. 2018 · 2018
Later among the works it cites.
More robust doubly robust off-policy evaluation. In ICML . 1447–1456
Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh. 2018 · 2018
Later among the works it cites.
Who should be treated? empirical welfare maximization methods for treatment choice
Toru Kitagawa and Aleksey Tetenov. 2018 · 2018
Later among the works it cites.
Competitive multi-agent inverse reinforcement learning with sub-optimal demonstrations
Xingyu Wang and Diego Klabjan. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Joshua Davidson, Christopher Archibald, and Michael Bowling. 2013 · 2013
Cited alongside, same era.
Using response functions to measure strategy strength. In AAAI . 630–636
Trevor Davis, Neil Burch, and Michael Bowling. 2014 · 2014
Cited alongside, same era.
Doubly robust policy evaluation and optimization
Miroslav Dudík, Dumitru Erhan, John Langford, Lihong Li, et al · 2014
Cited alongside, same era.
Offline policy evaluation across representations with applications to educational games.. In AAMAS . 1077–1084
Travis Mandel, Yun-En Liu, Sergey Levine, Emma Brunskill, and Zoran Popovic. 2014 · 2014
Cited alongside, same era.
Batch learning from logged bandit feedback through counterfactual risk minimization
Adith Swaminathan and Thorsten Joachims. 2015 · 2015
Cited alongside, same era.
Doubly Robust Off-policy Value Evaluation for Reinforcement Learning. In ICML . 652–661
Nan Jiang and Lihong Li. 2016 · 2016
Cited alongside, same era.
Safe, multi-agent, reinforcement learning for autonomous driving
Shai Shalev-Shwartz, Shaked Shammah, and Amnon Shashua. 2016 · 2016
Cited alongside, same era.
Kaiqing Zhang, Zhuoran Yang, Han Liu, Tong Zhang, and Tamer Başar. 2018 · 2018
Later among the works it cites.
Offline multi-action policy learning: Generalization and optimization
Zhengyuan Zhou, Susan Athey, and Stefan Wager. 2018 · 2018
Later among the works it cites.
Superhuman AI for multiplayer poker
Noam Brown and Tuomas Sandholm. 2019 · 2019
Later among the works it cites.
Low-Variance and Zero-Variance Baselines for Extensive-Form Games
Trevor Davis, Martin Schmid, and Michael Bowling. 2019 · 2019
Later among the works it cites.
Double reinforcement learning for efficient off-policy evaluation in markov decision processes
Nathan Kallus and Masatoshi Uehara. 2019a · 2019
Later among the works it cites.
Nathan Kallus and Masatoshi Uehara. 2019b · 2019
Later among the works it cites.
Variance reduction in monte carlo counterfactual regret minimization (VR-MCCFR) for extensive form games using baselines. In AAAI . 2157–2164
Martin Schmid, Neil Burch, Marc Lanctot, Matej Moravcik, Rudolf Kadlec, and Michael Bowling. 2019 · 2019
Later among the works it cites.
Towards Optimal Off-Policy Evaluation for Reinforcement Learning with Marginalized Importance Sampling. In NeurIPS . 9665–9675
Tengyang Xie, Yifei Ma, and Yu-Xiang Wang. 2019 · 2019
Later among the works it cites.
Multi-agent adversarial inverse reinforcement learning
Lantao Yu, Jiaming Song, and Stefano Ermon. 2019 · 2019
Later among the works it cites.
Multi-agent reinforcement learning: A selective overview of theories and algorithms
Kaiqing Zhang, Zhuoran Yang, and Tamer Başar. 2019a · 2019
Later among the works it cites.
Provable Self-Play Algorithms for Competitive Reinforcement Learning
Yu Bai and Chi Jin. 2020 · 2020
Closest in time.
Off-Policy Evaluation and Learning for External Validity under a Covariate Shift
Masahiro Kato, Masatoshi Uehara, and Shota Yasui. 2020 · 2020
Closest in time.