Fetching the paper…
Reading the bibliography…
Off-Policy Estimation (OPE) methods allow us to learn and evaluate decision-making policies from logged data.
Confounding and Simpson’s paradox
Steven A Julious and Mark A Mullee. 1994 · 1994
Earlier work this paper cites.
Off-policy Evaluation in Infinite-Horizon Reinforcement Learning with Latent Confounders. In Proceedings of The 24th International Conference on Artificial Intelligence and Statistics (Proceedings of Machine Learning Research, Vol. 130) , Arindam Banerjee and Kenji Fukumizu (Eds.). PMLR, 1999–2007
Andrew Bennett, Nathan Kallus, Lihong Li, and Ali Mousavi. 2021 · 2007
Earlier work this paper cites.
Causality
Judea Pearl. 2009 · 2009
Earlier work this paper cites.
Counterfactual Reasoning and Learning Systems: The Example of Computational Advertising
Léon Bottou, Jonas Peters, Joaquin Quiñonero-Candela, Denis X. Charles, D. Max Chickering, Elon Portugaly, Dipankar Ray, Patrice Simard, and Ed Snelson. 2013 · 2013
Earlier work this paper cites.
Testing for the Unconfoundedness Assumption Using an Instrumental Assumption
Xavier de Luna and Per Johansson. 2014 · 2013
Earlier work this paper cites.
An introduction to sensitivity analysis for unobserved confounding in nonexperimental prevention research
Weiwei Liu, S Janet Kuramoto, and Elizabeth A Stuart. 2013 · 2013
Earlier work this paper cites.
Monte Carlo theory, methods and examples
Art B. Owen. 2013 · 2013
Earlier work this paper cites.
Counterfactual Estimation and Optimization of Click Metrics in Search Engines: A Case Study. In Proceedings of the 24th International Conference on World Wide Web (WWW ’15 Companion) . ACM, 929–934
Lihong Li, Shunbao Chen, Jim Kleban, and Ankur Gupta. 2015 · 2015
Earlier work this paper cites.
The Self-Normalized Estimator for Counterfactual Learning. In Advances in Neural Information Processing Systems , C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett (Eds.), Vol. 28. Curran Associates, Inc
Adith Swaminathan and Thorsten Joachims. 2015 · 2015
Earlier work this paper cites.
Unbiased Learning to Rank with Unbiased Propensity Estimation. In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval (SIGIR ’18) . ACM, 385–394
Qingyao Ai, Keping Bi, Cheng Luo, Jiafeng Guo, and W. Bruce Croft. 2018 · 2018
Earlier work this paper cites.
Deep Learning with Logged Bandit Feedback. In International Conference on Learning Representations
Thorsten Joachims, Adith Swaminathan, and Maarten de Rijke. 2018 · 2018
Cited alongside, same era.
Unbiased Offline Recommender Evaluation for Missing-Not-at-Random Implicit Feedback. In Proceedings of the 12th ACM Conference on Recommender Systems (RecSys ’18) . ACM, 279–287
Longqi Yang, Yin Cui, Yuan Xuan, Chenyang Wang, Serge Belongie, and Deborah Estrin. 2018 · 2018
Cited alongside, same era.
Intervention Harvesting for Context-Dependent Examination-Bias Estimation. In Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR’19) . ACM, 825–834
Zhichong Fang, Aman Agarwal, and Thorsten Joachims. 2019 · 2019
Cited alongside, same era.
Importance Sampling Policy Evaluation with an Estimated Behavior Policy. In Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 97) , Kamalika Chaudhuri and Ruslan Salakhutdinov (Eds.). PMLR, 2605–2613
Josiah Hanna, Scott Niekum, and Peter Stone. 2019 · 2019
The Simpson’s Paradox in the Offline Evaluation of Recommendation Systems
Amir H. Jadidinejad, Craig Macdonald, and Iadh Ounis. 2021 · 2021
Later among the works it cites.
Counterfactual Learning and Evaluation for Recommender Systems: Foundations, Implementations, and Recent Advances. In Proc. of the 15th ACM Conference on Recommender Systems (RecSys ’21) . ACM, 828–830
Yuta Saito and Thorsten Joachims. 2021 · 2021
Later among the works it cites.
Provably Efficient Causal Reinforcement Learning with Confounded Observational Data. In Advances in Neural Information Processing Systems , M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan (Eds.), Vol. 34. Curran Associates, Inc., 21164–21175
Lingxiao Wang, Zhuoran Yang, and Zhaoran Wang. 2021 · 2021
Later among the works it cites.
Control Variate Diagnostics for Detecting Problems in Logged Bandit Feedback. In CONSEQUENCES+REVEAL Workshop at ACM RecSys ’22 (CONSEQUENCES+REVEAL ’22)
Ben London and Thorsten Joachims. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Revisiting Offline Evaluation for Implicit-Feedback Recommender Systems. In Proceedings of the 13th ACM Conference on Recommender Systems (RecSys ’19) . ACM, 596–600
Olivier Jeunen. 2019 · 2019
Cited alongside, same era.
On Sampled Metrics for Item Recommendation. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD ’20) . ACM, 1748–1757
Walid Krichene and Steffen Rendle. 2020 · 2020
Cited alongside, same era.
Off-policy Policy Evaluation For Sequential Decisions Under Unobserved Confounding. In Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 18819–18831
Hongseok Namkoong, Ramtin Keramati, Steve Yadlowsky, and Emma Brunskill. 2020 · 2020
Cited alongside, same era.
A Gentle Introduction to Recommendation as Counterfactual Policy Learning. In Proceedings of the 28th ACM Conference on User Modeling, Adaptation and Personalization (UMAP ’20) . ACM, 392–393
Flavian Vasile, David Rohde, Olivier Jeunen, and Amine Benhalloum. 2020 · 2020
Cited alongside, same era.
Causal Reinforcement Learning using Observational and Interventional Data
Maxime Gasse, Damien Grasset, Guillaume Gaudron, and Pierre-Yves Oudeyer. 2021 · 2021
Cited alongside, same era.
A Probabilistic Position Bias Model for Short-Video Recommendation Feeds. In Proceedings of the 17th ACM Conference on Recommender Systems (RecSys ’23) . ACM
Olivier Jeunen. 2023 · 2023
Closest in time.
Olivier Jeunen, Ivan Potapov, and Aleksei Ustimenko. 2023 · 2023
Closest in time.
Offline Policy Evaluation and Optimization under Confounding
Chinmaya Kausik, Yangyi Lu, Kevin Tan, Yixin Wang, and Ambuj Tewari. 2023 · 2023
Closest in time.
Take a Fresh Look at Recommender Systems from an Evaluation Standpoint. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’23) . ACM, 2629–2638
Aixin Sun. 2023 · 2023
Closest in time.
An Instrumental Variable Approach to Confounded Off-Policy Evaluation. In Proceedings of the 40th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 202) , Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (Eds.). PMLR, 38848–38880
Yang Xu, Jin Zhu, Chengchun Shi, Shikai Luo, and Rui Song. 2023 · 2023
Closest in time.