Fetching the paper…
Reading the bibliography…
Off-policy Learning to Rank (LTR) aims to optimize a ranker from data collected by a deployed logging policy.
Cumulated gain-based evaluation of ir techniques
Kalervo Järvelin and Jaana Kekäläinen · 2002
Earlier work this paper cites.
Optimizing search engines using clickthrough data
Thorsten Joachims · 2002
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Earlier work this paper cites.
Accurately interpreting clickthrough data as implicit feedback
Thorsten Joachims, Laura Granka, Bing Pan, Helene Hembrooke, and Geri Gay · 2005
Earlier work this paper cites.
Modeling result-list searching in the world wide web: The role of relevance topologies and trust bias
Maeve O’Brien and Mark Keane · 2006
Earlier work this paper cites.
An experimental comparison of click position-bias models
Nick Craswell, Onno Zoeter, Michael Taylor, and Bill Ramsey · 2008
Earlier work this paper cites.
A user browsing model to predict search engine click data from past observations
Georges E Dupret and Benjamin Piwowarski · 2008
Earlier work this paper cites.
Expected reciprocal rank for graded relevance
Olivier Chapelle, Donald Metlzer, Ya Zhang, and Pierre Grinspan · 2009
Earlier work this paper cites.
Learning to rank for information retrieval
Tie-Yan Liu et al · 2009
Earlier work this paper cites.
From ranknet to lambdarank to lambdamart: An overview
Chris J.C. Burges · 2010
Earlier work this paper cites.
Test collection based evaluation of information retrieval systems
Mark Sanderson et al · 2010
Earlier work this paper cites.
Yahoo! learning to rank challenge overview
Olivier Chapelle and Yi Chang · 2011
Earlier work this paper cites.
Doubly robust inference with missing data in survey sampling
Jae-kwang Kim and David Haziza · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Introducing LETOR 4.0 datasets
Tao Qin and Tie-Yan Liu · 2013
Earlier work this paper cites.
Click Models for Web Search
Aleksandr Chuklin, Ilya Markov, and Maarten de Rijke · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Counterfactual risk minimization: Learning from logged bandit feedback
Adith Swaminathan and Thorsten Joachims · 2015
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Cited alongside, same era.
Learning to rank with selection bias in personal search
Xuanhui Wang, Michael Bendersky, Donald Metzler, and Marc Najork · 2016
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Reinforcement learning to rank with markov decision process
Doubly robust estimator for ranking metrics with post-click conversions
Yuta Saito · 2020
Later among the works it cites.
Reinforcement learning to rank with pairwise policy gradient
Jun Xu, Zeng Wei, Long Xia, Yanyan Lan, Dawei Yin, Xueqi Cheng, and Ji-Rong Wen · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Y Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Later among the works it cites.
Offline reinforcement learning: fundamental barriers for value function approximation
Dylan J Foster, Akshay Krishnamurthy, David Simchi-Levi, and Yunzong Xu · 2021
Later among the works it cites.
Is pessimism provably efficient for offline RL?
Ying Jin, Zhuoran Yang, and Zhaoran Wang · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zeng Wei, Jun Xu, Yanyan Lan, Jiafeng Guo, and Xueqi Cheng · 2017
Cited alongside, same era.
Online learning to rank in stochastic click models
Masrour Zoghi, Tomas Tunys, Mohammad Ghavamzadeh, Branislav Kveton, Csaba Szepesvari, and Zheng Wen · 2017
Cited alongside, same era.
Unbiased learning to rank with unbiased propensity estimation
Qingyao Ai, Keping Bi, Cheng Luo, Jiafeng Guo, and W Bruce Croft · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abeel, and Sergey Levine · 2018
Cited alongside, same era.
Toprank: A practical algorithm for online stochastic ranking
Tor Lattimore, Branislav Kveton, Shuai Li, and Csaba Szepesvari · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Position bias estimation for unbiased learning to rank in personal search
Xuanhui Wang, Nadav Golbandi, Michael Bendersky, Donald Metzler, and Marc Najork · 2018
Cited alongside, same era.
Paria Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao, and Stuart Russell · 2021
Later among the works it cites.
Ultra: An unbiased learning to rank algorithm toolbox
Anh Tran, Tao Yang, and Qingyao Ai · 2021
Later among the works it cites.
Non-clicks mean irrelevant? propensity ratio scoring as a correction
Nan Wang, Zhen Qin, Xuanhui Wang, and Hongning Wang · 2021
Later among the works it cites.
Provable benefits of actor-critic methods for offline reinforcement learning
Andrea Zanette, Martin J Wainwright, and Emma Brunskill · 2021
Later among the works it cites.
Pessimistic off-policy optimization for learning to rank
Matej Cief, Branislav Kveton, and Michal Kompan · 2022
Later among the works it cites.
Doubly robust off-policy evaluation for ranking policies under the cascade behavior model
Haruka Kiyohara, Yuta Saito, Tatsuya Matsuhiro, Yusuke Narita, Nobuyuki Shimizu, and Yasuo Yamamoto · 2022
Later among the works it cites.
Pessimism for offline linear contextual bandits using ℓ p \ell_{p} confidence sets
Gene Li, Cong Ma, and Nathan Srebro · 2022
Later among the works it cites.
Pessimistic q-learning for offline reinforcement learning: Towards optimal sample complexity
Laixi Shi, Gen Li, Yuting Wei, Yuxin Chen, and Yuejie Chi · 2022
Later among the works it cites.
Provably efficient offline reinforcement learning with trajectory-wise reward
Tengyu Xu and Yingbin Liang · 2022
Later among the works it cites.
Reinforcement online learning to rank with unbiased reward shaping
Shengyao Zhuang, Zhihao Qiao, and Guido Zuccon · 2022
Later among the works it cites.
Doubly robust estimation for correcting position bias in click feedback for unbiased learning to rank
Harrie Oosterhuis · 2023
Closest in time.
Cascade model-based propensity estimation for counterfactual learning to rank
Ali Vardasbi, Maarten de Rijke, and Ilya Markov · 2092
Closest in time.