Fetching the paper…
Reading the bibliography…
Off-policy learning is a framework for optimizing policies without deploying them, using data collected by another policy.
The Optimizer’s Curse: Skepticism and Postdecision Surprise in Decision Analysis
J. E. Smith and R. L. Winkler · 1909
Earlier work this paper cites.
DCM Bandits: Learning to Rank with Multiple Clicks
S. Katariya, B. Kveton, C. Szepesvari, and Z. Wen · 1938
Earlier work this paper cites.
Cascading Bandits: Learning to Rank in the Cascade Model
B. Kveton, C. Szepesvari, Z. Wen, and A. Ashkan · 1938
Earlier work this paper cites.
A Generalization of Sampling Without Replacement from a Finite Universe
D. G. Horvitz and D. J. Thompson · 1952
Earlier work this paper cites.
Estimation of Regression Coefficients When Some Regressors are not Always Observed
J. M. Robins, A. Rotnitzky, and L. P. Zhao · 1994
Earlier work this paper cites.
Optimizing search engines using clickthrough data
T. Joachims · 2002
Earlier work this paper cites.
Pattern recognition and machine learning
C. M. Bishop · 2006
Earlier work this paper cites.
Evaluating the accuracy of implicit feedback from clicks and query reformulations in Web search
T. Joachims, L. Granka, B. Pan, H. Hembrooke, F. Radlinski, and G. Gay · 2007
Earlier work this paper cites.
Predicting clicks: estimating the click-through rate for new ads
M. Richardson, E. Dominowska, and R. Ragno · 2007
Earlier work this paper cites.
An experimental comparison of click position-bias models
N. Craswell, O. Zoeter, M. Taylor, and B. Ramsey · 2008
Earlier work this paper cites.
Truncated Importance Sampling
E. L. Ionides · 2008
Earlier work this paper cites.
Expected reciprocal rank for graded relevance
O. Chapelle, D. Metlzer, Y. Zhang, and P. Grinspan · 2009
Earlier work this paper cites.
Efficient multiple-click models in web search
F. Guo, C. Liu, and Y. M. Wang · 2009
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
L. Li, W. Chu, J. Langford, and R. E. Schapire · 2010
Earlier work this paper cites.
Learning from Logged Implicit Exploration Data
A. Strehl, J. Langford, L. Li, and S. M. Kakade · 2010
Cited alongside, same era.
C14 - Yahoo! Learning to Rank Challenge, 2010
Yahoo! · 2010
Cited alongside, same era.
Alternating least squares for personalized ranking
G. Takács and D. Tikk · 2012
Cited alongside, same era.
Counterfactual Reasoning and Learning Systems: The Example of Computational Advertising
L. Bottou, J. Peters, J. Quiñonero-Candela, D. X. Charles, D. M. Chickering, E. Portugaly, D. Ray, P. Simard, and E. Snelson · 2013
Cited alongside, same era.
Fidelity, Soundness, and Efficiency of Interleaved Comparison Methods
K. Hofmann, S. Whiteson, and M. D. Rijke · 2013
Cited alongside, same era.
Introducing LETOR 4.0 Datasets, June 2013
T. Qin and T.-Y. Liu · 2013
Off-policy evaluation for slate recommendation
A. Swaminathan, A. Krishnamurthy, A. Agarwal, M. Dudik, J. Langford, D. Jose, and I. Zitouni · 2017
Later among the works it cites.
Offline A/B Testing for Recommender Systems
A. Gilotte, C. Calauzènes, T. Nedelec, A. Abraham, and S. Dollé · 2018
Later among the works it cites.
Deep Learning with Logged Bandit Feedback
T. Joachims, A. Swaminathan, and M. d. Rijke · 2018
Later among the works it cites.
Offline Evaluation of Ranking Policies with Click Models
S. Li, Y. Abbasi-Yadkori, B. Kveton, S. Muthukrishnan, V. Vinay, and Z. Wen · 2018
Later among the works it cites.
Explore, exploit, and explain: personalizing explainable recommendations with bandits
J. McInerney, B. Lacker, S. Hansen, K. Higley, H. Bouchard, A. Gruson, and R. Mehrotra · 2018
Later among the works it cites.
High-Dimensional Probability: An Introduction with Applications in Data Science
R. Vershynin · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Yandex Personalized Web Search Challenge, 2013
Yandex · 2013
Cited alongside, same era.
Doubly Robust Policy Evaluation and Optimization
M. Dudik, D. Erhan, J. Langford, and L. Li · 2014
Cited alongside, same era.
Practical Lessons from Predicting Clicks on Ads at Facebook
X. He, J. Pan, O. Jin, T. Xu, B. Liu, T. Xu, Y. Shi, A. Atallah, R. Herbrich, S. Bowers, and J. Q. Candela · 2014
Cited alongside, same era.
Click Models for Web Search
A. Chuklin, I. Markov, and M. De Rijke · 2015
Cited alongside, same era.
The Self-Normalized Estimator for Counterfactual Learning
A. Swaminathan and T. Joachims · 2015
Cited alongside, same era.
Fast Ranking with Additive Ensembles of Oblivious and Non-Oblivious Regression Trees
D. Dato, C. Lucchese, F. M. Nardini, S. Orlando, R. Perego, N. Tonellotto, and R. Venturini · 2017
Cited alongside, same era.
Later among the works it cites.
Joint Policy-Value Learning for Recommendation
O. Jeunen, D. Rohde, F. Vasile, and M. Bompaire · 2020
Later among the works it cites.
Non-Stationary Off-Policy Optimization
J. Hong, B. Kveton, M. Zaheer, Y. Chow, and A. Ahmed · 2021
Later among the works it cites.
Pessimistic Reward Models for Off-Policy Learning in Recommendation
O. Jeunen and B. Goethals · 2021
Later among the works it cites.
Is Pessimism Provably Efficient for Offline RL?
Y. Jin, Z. Yang, and Z. Wang · 2021
Later among the works it cites.
Bellman-consistent Pessimism for Offline Reinforcement Learning
T. Xie, C.-A. Cheng, N. Jiang, P. Mineiro, and A. Agarwal · 2021
Later among the works it cites.
Doubly Robust Off-Policy Evaluation for Ranking Policies under the Cascade Behavior Model
H. Kiyohara, Y. Saito, T. Matsuhiro, Y. Narita, N. Shimizu, and Y. Yamamoto · 2022
Closest in time.
Multi-Task Off-Policy Learning from Bandit Feedback
J. Hong, B. Kveton, M. Zaheer, S. Katariya, and M. Ghavamzadeh · 2023
Closest in time.