Fetching the paper…
Reading the bibliography…
Learning-to-rank (LTR) has become a key technology in E-commerce applications.
Introduction to reinforcement learning
RS Sutton and AG Barto · 1998
Earlier work this paper cites.
Optimizing search engines using clickthrough data
Thorsten Joachims · 2002
Earlier work this paper cites.
Learning to rank using gradient descent
Christopher J. C. Burges, Tal Shaked, Erin Renshaw, Ari Lazier, Matt Deeds, Nicole Hamilton, and Gregory N. Hullender · 2005
Earlier work this paper cites.
Learning to rank with nonsmooth cost functions
Christopher J. C. Burges, Robert Ragno, and Quoc Viet Le · 2006
Earlier work this paper cites.
Being accurate is not enough: how accuracy metrics have hurt recommender systems
Sean M McNee, John Riedl, and Joseph A Konstan · 2006
Earlier work this paper cites.
Learning to rank: from pairwise approach to listwise approach
Zhe Cao, Tao Qin, Tie-Yan Liu, Ming-Feng Tsai, and Hang Li · 2007
Earlier work this paper cites.
Mcrank: Learning to rank using multiple classification and gradient boosting
Ping Li, Christopher J. C. Burges, and Qiang Wu · 2007
Earlier work this paper cites.
Statistical analysis of bayes optimal subset ranking
David Cossock and Tong Zhang · 2008
Earlier work this paper cites.
Softrank: optimizing non-smooth rank metrics
Michael Taylor, John Guiver, Stephen Robertson, and Tom Minka · 2008
Earlier work this paper cites.
Listwise approach to learning to rank: theory and algorithm
Fen Xia, Tie-Yan Liu, Jue Wang, Wensheng Zhang, and Hang Li · 2008
Earlier work this paper cites.
From ranknet to lambdarank to lambdamart: An overview
Christopher JC Burges · 2010
Earlier work this paper cites.
Efficient reductions for imitation learning
Stéphane Ross and Drew Bagnell · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Kiri Wagstaff · 2012
Earlier work this paper cites.
A comparative analysis of offline and online evaluations and discussion of research paper recommender system evaluation
Joeran Beel, Marcel Genzmehr, Stefan Langer, Andreas Nürnberger, and Bela Gipp · 2013
Earlier work this paper cites.
Generative adversarial nets
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio · 2014
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael I. Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Pointer networks
Oriol Vinyals, Meire Fortunato, and Navdeep Jaitly · 2015
Cited alongside, same era.
Guided cost learning: Deep inverse optimal control via policy optimization
Chelsea Finn, Sergey Levine, and Pieter Abbeel · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Cited alongside, same era.
Contrasting offline and online results when evaluating recommendation algorithms
Marco Rossetti, Fabio Stella, and Markus Zanker · 2016
Cited alongside, same era.
Deep interest network for click-through rate prediction
Guorui Zhou, Xiaoqiang Zhu, Chenru Song, Ying Fan, Han Zhu, Xiao Ma, Yanghui Yan, Junqi Jin, Han Li, and Kun Gai · 2018
Later among the works it cites.
Globally optimized mutual influence aware ranking in e-commerce search
Tao Zhuang, Wenwu Ou, and Zhirong Wang · 2018
Later among the works it cites.
A general framework for counterfactual learning-to-rank
Aman Agarwal, Kenta Takatsu, Ivan Zaitsev, and Thorsten Joachims · 2019
Later among the works it cites.
Learning groupwise multivariate scoring functions using deep neural networks
Qingyao Ai, Xuanhui Wang, Sebastian Bruch, Nadav Golbandi, Michael Bendersky, and Marc Najork · 2019
Later among the works it cites.
Top-k off-policy correction for a REINFORCE recommender system
Minmin Chen, Alex Beutel, Paul Covington, Sagar Jain, Francois Belletti, and Ed H. Chi · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
One-shot imitation learning
Yan Duan, Marcin Andrychowicz, Bradly Stadie, OpenAI Jonathan Ho, Jonas Schneider, Ilya Sutskever, Pieter Abbeel, and Wojciech Zaremba · 2017
Cited alongside, same era.
Unbiased learning-to-rank with biased feedback
Thorsten Joachims, Adith Swaminathan, and Tobias Schnabel · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Learning a deep listwise context model for ranking refinement
Qingyao Ai, Keping Bi, Jiafeng Guo, and W Bruce Croft · 2018
Cited alongside, same era.
Seq2slate: Re-ranking and slate optimization with rnns
Irwan Bello, Sayali Kulkarni, Sagar Jain, Craig Boutilier, Ed Huai-hsin Chi, Elad Eban, Xiyang Luo, Alan Mackey, and Ofer Meshi · 2018
Cited alongside, same era.
Entire space multi-task model: An effective approach for estimating post-click conversion rate
Xiao Ma, Liqin Zhao, Guan Huang, Zhi Wang, Zelin Hu, Xiaoqiang Zhu, and Kun Gai · 2018
Cited alongside, same era.
Are we really making much progress? a worrying analysis of recent neural recommendation approaches
Maurizio Ferrari Dacrema, Paolo Cremonesi, and Dietmar Jannach · 2019
Later among the works it cites.
Exact-k recommendation via maximal clique optimization
Yu Gong, Yu Zhu, Lu Duan, Qingwen Liu, Ziyu Guan, Fei Sun, Wenwu Ou, and Kenny Q. Zhu · 2019
Later among the works it cites.
Recsim: A configurable simulation platform for recommender systems
Eugene Ie, Chih-wei Hsu, Martin Mladenov, Vihan Jain, Sanmit Narvekar, Jing Wang, Rui Wu, and Craig Boutilier · 2019
Later among the works it cites.
Slateq: A tractable decomposition for reinforcement learning with recommendation sets
Eugene Ie, Vihan Jain, Jing Wang, Sanmit Narvekar, Ritesh Agarwal, Rui Wu, Heng-Tze Cheng, Tushar Chandra, and Craig Boutilier · 2019
Later among the works it cites.
Beyond greedy ranking: Slate optimization via list-cvae
Ray Jiang, Sven Gowal, Yuqiu Qian, Timothy A. Mann, and Danilo J. Rezende · 2019
Later among the works it cites.
Buy 4 reinforce samples, get a baseline for free!
Wouter Kool, Herke van Hoof, and Max Welling · 2019
Later among the works it cites.
Personalized re-ranking for recommendation
Changhua Pei, Yi Zhang, Yongfeng Zhang, Fei Sun, Xiao Lin, Hanxiao Sun, Jian Wu, Peng Jiang, Junfeng Ge, Wenwu Ou, et al · 2019
Later among the works it cites.
Virtual-taobao: Virtualizing real-world online retail environment for reinforcement learning
Jing-Cheng Shi, Yang Yu, Qing Da, Shi-Yong Chen, and An-Xiang Zeng · 2019
Later among the works it cites.
Sequential evaluation and generation framework for combinatorial recommender system
Fan Wang, Xiaomin Fang, Lihang Liu, Yaxue Chen, Jiucheng Tao, Zhiming Peng, Cihang Jin, and Hao Tian · 2019
Later among the works it cites.
Context-aware ranking by constructing a virtual environment for reinforcement learning
Junqi Zhang, Jiaxin Mao, Yiqun Liu, Ruizhe Zhang, Min Zhang, Shaoping Ma, Jun Xu, and Qi Tian · 2019
Later among the works it cites.