Fetching the paper…
Reading the bibliography…
Modern recommendation systems ought to benefit by probing for and learning from delayed feedback.
Eugene Ie, Vihan Jain, Jing Wang, Sanmit Narvekar, Ritesh Agarwal, Rui Wu, Heng-Tze Cheng, Tushar Chandra, and Craig Boutilier. 2019 · 1905
Earlier work this paper cites.
On the Likelihood That One Unknown Probability Exceeds Another in View of the Evidence of Two Samples
William R Thompson. 1933 · 1933
Earlier work this paper cites.
The Multi-Armed Bandit Problem: Decomposition and Computation
Michael N Katehakis and Arthur F Veinott Jr. 1987 · 1987
Earlier work this paper cites.
Finite-Time Analysis of the Multiarmed Bandit Problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer. 2002 · 2002
Earlier work this paper cites.
On the Sample Complexity of Reinforcement Learning
Sham Machandranath Kakade. 2003 · 2003
Earlier work this paper cites.
Reinforcement Learning Architecture for Web Recommendations. In International Conference on Information Technology: Coding and Computing, 2004. Proceedings. ITCC 2004. , Vol. 1. IEEE, 398–402
Nick Golovin and Erhard Rahm. 2004 · 2004
Earlier work this paper cites.
Integrating AHP and Data Mining for Product Recommendation Based on Customer Lifetime Vsalue
Duen-Ren Liu and Ya-Yueh Shih. 2005 · 2005
Earlier work this paper cites.
New Recommendation System Using Reinforcement Learning
Pornthep Rojanavasu, Phaitoon Srinil, and Ouen Pinngern. 2005 · 2005
Earlier work this paper cites.
Epoch-Greedy Algorithm for Multi-Armed Bandits with Side Information
John Langford and Tong Zhang. 2007 · 2007
Earlier work this paper cites.
Collaborative Filtering Recommender Systems
J Ben Schafer, Dan Frankowski, Jon Herlocker, and Shilad Sen. 2007 · 2007
Earlier work this paper cites.
Recommendation Method for Improving Customer Lifetime Value
Tomoharu Iwata, Kazumi Saito, and Takeshi Yamada. 2008 · 2008
Earlier work this paper cites.
Product Recommendation Approaches: Collaborative Filtering via Customer Lifetime Value and Customer Demands
Ya-Yueh Shih and Duen-Ren Liu. 2008 · 2008
Earlier work this paper cites.
Understanding the Difficulty of Training Deep Feedforward Neural Networks. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics . 249–256
Xavier Glorot and Yoshua Bengio. 2010 · 2010
Earlier work this paper cites.
A Contextual-Bandit Approach to Personalized News Article Recommendation. In Proceedings of the 19th International Conference on World Wide Web . 661–670
Lihong Li, Wei Chu, John Langford, and Robert E Schapire. 2010 · 2010
Earlier work this paper cites.
An Empirical Evaluation of Thompson Sampling. In Advances in Neural Information Processing Systems . 2249–2257
Olivier Chapelle and Lihong Li. 2011 · 2011
Earlier work this paper cites.
The KL-UCB Algorithm for Bounded Stochastic Bandits and Beyond. In Proceedings of the 24th annual conference on learning theory . JMLR Workshop and Conference Proceedings, 359–376
Aurélien Garivier and Olivier Cappé. 2011 · 2011
Earlier work this paper cites.
Personalized Pricing Recommender System: Multi-Stage Epsilon-Greedy Approach. In Proceedings of the 2nd International Workshop on Information Heterogeneity and Fusion in Recommender Systems . 57–64
Toshihiro Kamishima and Shotaro Akaho. 2011 · 2011
Earlier work this paper cites.
Unbiased Offline Evaluation of Contextual-Bandit-Based News Article Recommendation Algorithms. In Proceedings of the Fourth ACM International Conference on Web Search and Data Mining . ACM, 297–306
Lihong Li, Wei Chu, John Langford, and Xuanhui Wang. 2011 · 2011
Earlier work this paper cites.
A Contextual-Bandit Algorithm for Mobile Context-Aware Recommender System. In International Conference on Neural Information Processing . Springer, 324–331
Djallel Bouneffouf, Amel Bouzeghoub, and Alda Lopes Gançarski. 2012 · 2012
Earlier work this paper cites.
An Overview of Classification Algorithms for Imbalanced Datasets
Vaishali Ganganwar. 2012 · 2012
Earlier work this paper cites.
Online Learning under Delayed Feedback. In International Conference on Machine Learning . 1453–1461
Pooria Joulani, Andras Gyorgy, and Csaba Szepesvári. 2013 · 2013
Earlier work this paper cites.
Learning to Optimize via Information-Directed Sampling. In Advances in Neural Information Processing Systems . 1583–1591
Daniel Russo and Benjamin Van Roy. 2014 · 2014
Earlier work this paper cites.
Human-Level Control through Deep Reinforcement Learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
A UCB-Like Strategy of Collaborative Filtering. In Asian Conference on Machine Learning . 315–329
Atsuyoshi Nakamura. 2015 · 2015
Cited alongside, same era.
Personalized Ad Recommendation Systems for Life-Time Value Optimization with Guarantees. In Twenty-Fourth International Joint Conference on Artificial Intelligence
Georgios Theocharous, Philip S Thomas, and Mohammad Ghavamzadeh. 2015 · 2015
Cited alongside, same era.
Online Recommender Systems–How Does a Website Know What I Want?
Stephanie Blanda. 2016 · 2016
Cited alongside, same era.
Wide & Deep Learning for Recommender Systems. In Proceedings of the 1st workshop on deep learning for recommender systems . 7–10
Heng-Tze Cheng, Levent Koc, Jeremiah Harmsen, Tal Shaked, Tushar Chandra, Hrishi Aradhye, Glen Anderson, Greg Corrado, Wei Chai, Mustafa Ispir, et al · 2016
Addressing Delayed Feedback for Continuous Training with Neural Networks in CTR Prediction. In Proceedings of the 13th ACM Conference on Recommender Systems . 187–195
Sofia Ira Ktena, Alykhan Tejani, Lucas Theis, Pranay Kumar Myana, Deepak Dilipkumar, Ferenc Huszár, Steven Yoo, and Wenzhe Shi. 2019 · 2019
Later among the works it cites.
Recommendation System-Based Upper Confidence Bound for Online Advertising. In REVEAL 2019
Nhan Nguyen-Thanh, Dana Marinca, Kinda Khawam, David Rohde, Flavian Vasile, Elena Lohan, Steven Martin, and Dominique Quadri. 2019 · 2019
Later among the works it cites.
Deep Exploration via Randomized Value Functions
Ian Osband, Benjamin Van Roy, Daniel J Russo, and Zheng Wen. 2019 · 2019
Later among the works it cites.
Virtual-Taobao: Virtualizing Real-World Online Retail Environment for Reinforcement Learning. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 33. 4902–4909
Jing-Cheng Shi, Yang Yu, Qing Da, Shi-Yong Chen, and An-Xiang Zeng. 2019 · 2019
Later among the works it cites.
A Reinforcement Learning Approach to Personalized Learning Recommendation Systems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep Exploration via Bootstrapped DQN. In Advances in Neural information processing systems
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy. 2016 · 2016
Cited alongside, same era.
Fairness in Machine Learning
Solon Barocas, Moritz Hardt, and Arvind Narayanan. 2017 · 2017
Cited alongside, same era.
Spectrally-Normalized Margin Bounds for Neural Networks. In Advances in Neural Information Processing Systems . 6240–6249
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky. 2017 · 2017
Cited alongside, same era.
Playlist Recommendation Based on Reinforcement Learning. In International Conference on Intelligence Science . Springer, 172–182
Binbin Hu, Chuan Shi, and Jian Liu. 2017 · 2017
Cited alongside, same era.
Ensemble Sampling. In Advances in Neural Information Processing Systems . 3258–3266
Xiuyuan Lu and Benjamin Van Roy. 2017 · 2017
Cited alongside, same era.
Mastering the Game of Go without Human Knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
Understanding Deep Learning Requires Rethinking Generalization Int. In Conf. on Learning Representations
C Zhang, S Bengio, M Hardt, B Recht, and O Vinyals. 2017 · 2017
Cited alongside, same era.
Xueying Tang, Yunxiao Chen, Xiaoou Li, Jingchen Liu, and Zhiliang Ying. 2019 · 2019
Later among the works it cites.
Neural News Recommendation with Multi-Head Self-Attention. In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP) . 6389–6394
Chuhan Wu, Fangzhao Wu, Suyu Ge, Tao Qi, Yongfeng Huang, and Xing Xie. 2019 · 2019
Later among the works it cites.
Deep Interest Evolution Network for Click-Through Rate Prediction. In Proceedings of the AAAI conference on artificial intelligence , Vol. 33. 5941–5948
Guorui Zhou, Na Mou, Ying Fan, Qi Pi, Weijie Bian, Chang Zhou, Xiaoqiang Zhu, and Kun Gai. 2019 · 2019
Later among the works it cites.
Generative Adversarial Networks
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2020 · 2020
Later among the works it cites.
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Later among the works it cites.
Neural Thompson Sampling. In International Conference on Learning Representations
Weitong Zhang, Dongruo Zhou, Lihong Li, and Quanquan Gu. 2020 · 2020
Later among the works it cites.
Neural Contextual Bandits with UCB-Based Exploration. In International Conference on Machine Learning . PMLR, 11492–11502
Dongruo Zhou, Lihong Li, and Quanquan Gu. 2020 · 2020
Later among the works it cites.
Values of User Exploration in Recommender Systems. In Fifteenth ACM Conference on Recommender Systems . 85–95
Minmin Chen, Yuyan Wang, Can Xu, Ya Le, Mohit Sharma, Lee Richardson, Su-Lin Wu, and Ed Chi. 2021 · 2021
Closest in time.
Ian Osband, Zheng Wen, Mohammad Asghari, Morteza Ibrahimi, Xiyuan Lu, and Benjamin Van Roy. 2021 · 2021
Closest in time.
RL4RS: A Real-World Benchmark for Reinforcement Learning based Recommender System
Kai Wang, Zhene Zou, Qilin Deng, Yue Shang, Minghao Zhao, Runze Wu, Xudong Shen, Tangjie Lyu, and Changjie Fan. 2021 · 2021
Closest in time.
Neural Contextual Bandits with Deep Representation and Shallow Exploration. In International Conference on Learning Representations
Pan Xu, Zheng Wen, Handong Zhao, and Quanquan Gu. 2021 · 2021
Closest in time.
Off-Policy Actor-critic for Recommender Systems. In Proceedings of the 16th ACM Conference on Recommender Systems . 338–349
Minmin Chen, Can Xu, Vince Gatto, Devanshu Jain, Aviral Kumar, and Ed Chi. 2022 · 2022
Closest in time.
The Neural Testbed: Evaluating Joint Predictions. In Advances in Neural Information Processing Systems
Ian Osband, Zhegn Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla, Xiuyuan Lu, Morteza Ibrahimi, Dieterich Lawson, Botao Hao, Brendan O’Donoghue, and Benjamin Van Roy. 2022 · 2022
Closest in time.
An Analysis of Ensemble Sampling. In Advances in Neural Information Processing Systems
Chao Qin, Zheng Wen, Xiuyuan Lu, and Benjamin Van Roy. 2022 · 2022
Closest in time.
Evaluating Online Bandit Exploration In Large-Scale Recommender System. In KDD-23 Workshop on Multi-Armed Bandits and Reinforcement Learning: Advancing Decision Making in E-Commerce and Beyond
Hongbo Guo, Ruben Naeff, Alex Nikulkov, and Zheqing Zhu. 2023 · 2023
Closest in time.
Optimizing Long-term Value for Auction-Based Recommender Systems via On-Policy Reinforcement Learning. In Proceedings of the 17th ACM Conference on Recommender Systems
Ruiyang Xu, Jalaj Bhandari, Dmytro Korenkevych, Fan Liu, Yuchen He, Alex Nikulkov, and Zheqing Zhu. 2023 · 2023
Closest in time.
Scalable Neural Contextual Bandit for Recommender Systems
Zheqing Zhu and Benjamin Van Roy. 2023 · 2023
Closest in time.