Fetching the paper…
Reading the bibliography…
In e-commerce platforms such as Amazon and TaoBao, ranking items in a search session is a typical multi-step decision-making problem.
Learning from delayed rewards
C.J.C.H. Watkins. 1989 · 1989
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams. 1992 · 1992
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R.S. Sutton and A.G. Barto. 1998 · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation. In Advances in Neural Information Processing Systems (NIPS’00) . 1057–1063
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour. 2000 · 2000
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer. 2002 · 2002
Earlier work this paper cites.
R-MAX - A General Polynomial Time Algorithm for Near-Optimal Reinforcement Learning
Ronen I. Brafman and Moshe Tennenholtz. 2002 · 2002
Earlier work this paper cites.
Optimizing search engines using clickthrough data. In Proceedings of the eighth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD’02) . ACM, 133–142
Thorsten Joachims. 2002 · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh. 2002 · 2002
Earlier work this paper cites.
Discriminative models for information retrieval. In Proceedings of the 27th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR’04) . ACM, 64–71
Ramesh Nallapati. 2004 · 2004
Earlier work this paper cites.
Learning to rank using gradient descent. In Proceedings of the 22nd International Conference on Machine Learning . 89–96
Chris Burges, Tal Shaked, Erin Renshaw, Ari Lazier, Matt Deeds, Nicole Hamilton, and Greg Hullender. 2005 · 2005
Earlier work this paper cites.
Adapting ranking SVM to document retrieval. In Proceedings of the 29th Annual International Conference on Research and Development in Information Retrieval (SIGIR’06) . 186–193
Yunbo Cao, Jun Xu, Tie-Yan Liu, Hang Li, Yalou Huang, and Hsiao-Wuen Hon. 2006 · 2006
Earlier work this paper cites.
Learning to rank: from pairwise approach to listwise approach. In Proceedings of the 24th International Conference on Machine Learning (ICML’07) . ACM, 129–136
Zhe Cao, Tao Qin, Tie-Yan Liu, Ming-Feng Tsai, and Hang Li. 2007 · 2007
Cited alongside, same era.
The epoch-greedy algorithm for multi-armed bandits with side information. In Advances in neural information processing systems . 817–824
John Langford and Tong Zhang. 2008 · 2008
Cited alongside, same era.
Mcrank: Learning to rank using multiple classification and gradient boosting. In Advances in Neural Information Processing Systems (NIPS’08) . 897–904
Ping Li, Qiang Wu, and Christopher J Burges. 2008 · 2008
Cited alongside, same era.
Learning diverse rankings with multi-armed bandits. In Proceedings of the 25th international conference on Machine learning . ACM, 784–791
Filip Radlinski, Robert Kleinberg, and Thorsten Joachims. 2008 · 2008
Cited alongside, same era.
Learning to rank for information retrieval
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2015 · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Later among the works it cites.
Trust region policy optimization. In Proceedings of the 32nd International Conference on Machine Learning (ICML’15) . 1889–1897
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz. 2015 · 2015
Later among the works it cites.
Multiple-play bandits in the position-based model. In Advances in Neural Information Processing Systems (NIPS’16) . 1597–1605
Paul Lagrée, Claire Vernade, and Olivier Cappe. 2016 · 2016
Later among the works it cites.
Contextual combinatorial cascading bandits. In International Conference on Machine Learning (ICML’16) . 1245–1253
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tie-Yan Liu et al · 2009
Cited alongside, same era.
Interactively optimizing information retrieval systems as a dueling bandits problem. In Proceedings of the 26th Annual International Conference on Machine Learning (ICML’09) . ACM, 1201–1208
Yisong Yue and Thorsten Joachims. 2009 · 2009
Cited alongside, same era.
Toward off-policy learning control with function approximation. In Proceedings of the 27th International Conference on Machine Learning . 719–726
Hamid R. Maei, Csaba Szepesvári, Shalabh Bhatnagar, and Richard S. Sutton. 2010 · 2010
Cited alongside, same era.
Balancing exploration and exploitation in listwise and pairwise online learning to rank for information retrieval
Katja Hofmann, Shimon Whiteson, and Maarten de Rijke. 2013 · 2013
Cited alongside, same era.
Ranked bandits in metric spaces: learning diverse rankings over large document collections
Aleksandrs Slivkins, Filip Radlinski, and Sreenivas Gollapudi. 2013 · 2013
Cited alongside, same era.
Deterministic policy gradient algorithms. In Proceedings of the 31st International Conference on Machine Learning (ICML’14) . 387–395
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller. 2014 · 2014
Cited alongside, same era.
Cascading bandits: Learning to rank in the cascade model. In Proceedings of the 32nd International Conference on Machine Learning (ICML-15) . 767–776
Branislav Kveton, Csaba Szepesvari, Zheng Wen, and Azin Ashkan. 2015a
Cited in the paper.
Combinatorial cascading bandits. In Advances in Neural Information Processing Systems (NIPS’15) . 1450–1458
Branislav Kveton, Zheng Wen, Azin Ashkan, and Csaba Szepesvari. 2015b
Cited in the paper.
Shuai Li, Baoxiang Wang, Shengyu Zhang, and Wei Chen. 2016 · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Later among the works it cites.
Cascading bandits for large-scale recommendation problems. In Proceedings of the Thirty-Second Conference on Uncertainty in Artificial Intelligence (UAI’16) . 835–844
Shi Zong, Hao Ni, Kenny Sung, Nan Rosemary Ke, Zheng Wen, and Branislav Kveton. 2016 · 2016
Later among the works it cites.
Joe Tsai Looks Beyond Alibaba’s RMB 3 Trillion Milestone
Alizila. 2017 · 2017
Later among the works it cites.
Stochastic Rank-1 Bandits. In Artificial Intelligence and Statistics . 392–401
Sumeet Katariya, Branislav Kveton, Csaba Szepesvari, Claire Vernade, and Zheng Wen. 2017 · 2017
Later among the works it cites.
Online Learning to Rank in Stochastic Click Models. In International Conference on Machine Learning . 4199–4208
Masrour Zoghi, Tomas Tunys, Mohammad Ghavamzadeh, Branislav Kveton, Csaba Szepesvari, and Zheng Wen. 2017 · 2017
Later among the works it cites.