Fetching the paper…
Reading the bibliography…
With the recent prevalence of Reinforcement Learning (RL), there have been tremendous interests in utilizing RL for online advertising in recommendation platforms (e.g., e-commerce and news feed sites).
Toward Simulating Environments in Reinforcement Learning Based Recommendations
Zhao, X.; Xia, L.; Ding, Z.; Yin, D.; and Tang, J. 2019a · 1906
Earlier work this paper cites.
Learning to Collaborate: Multi-Scenario Ranking via Multi-Agent Reinforcement Learning
Feng, J.; Li, H.; Huang, M.; Liu, S.; Ou, W.; Wang, Z.; and Zhu, X. 2018 · 1948
Earlier work this paper cites.
Attacking Black-box Recommendations via Copying Cross-domain User Profiles
Fan, W.; Derr, T.; Zhao, X.; Ma, Y.; Liu, H.; Wang, J.; Tang, J.; and Li, Q. 2020 · 2005
Earlier work this paper cites.
A unified optimization framework for auction and guaranteed delivery in online advertising
Salomatin, K.; Liu, T.-Y.; and Yang, Y. 2012 · 2009
Earlier work this paper cites.
Degris, T.; White, M.; and Sutton, R. S. 2012 · 2012
Earlier work this paper cites.
Dynamic programming
Bellman, R. 2013 · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Graves, A.; Antonoglou, I.; Wierstra, D.; and Riedmiller, M. 2013 · 2013
Earlier work this paper cites.
Automatic ad format selection via contextual bandits
Tang, L.; Rosales, R.; Singh, A.; and Agarwal, D. 2013 · 2013
Earlier work this paper cites.
Estimation bias in multi-armed bandit algorithms for search advertising
Xu, M.; Qin, T.; and Liu, T.-Y. 2013 · 2013
Earlier work this paper cites.
Adaptive keywords extraction with contextual bandits for advertising on parked domains
Yuan, S.; Wang, J.; and van der Meer, M. 2013 · 2013
Earlier work this paper cites.
Convolutional Neural Networks for Sentence Classification
Kim, Y. 2014 · 2014
Earlier work this paper cites.
Deep reinforcement learning in large discrete action spaces
Dulac-Arnold, G.; Evans, R.; van Hasselt, H.; Sunehag, P.; Lillicrap, T.; Hunt, J.; Mann, T.; Weber, T.; Degris, T.; and Coppin, B. 2015 · 2015
Earlier work this paper cites.
Session-based recommendations with recurrent neural networks
Hidasi, B.; Karatzoglou, A.; Baltrunas, L.; and Tikk, D. 2015 · 2015
Cited alongside, same era.
Dueling Network Architectures for Deep Reinforcement Learning
Wang, Z.; Freitas, N. D.; and Lanctot, M. 2015 · 2015
Cited alongside, same era.
Brockman, G.; Cheung, V.; Pettersson, L.; Schneider, J.; Schulman, J.; Tang, J.; and Zaremba, W. 2016 · 2016
Cited alongside, same era.
Wide & deep learning for recommender systems
Cheng, H.-T.; Koc, L.; Harmsen, J.; Shaked, T.; Chandra, T.; Aradhye, H.; Anderson, G.; Corrado, G.; Chai, W.; Ispir, M.; et al. 2016 · 2016
Cited alongside, same era.
Efficient Delivery Policy to Minimize User Traffic Consumption in Guaranteed Advertising
Jia, Z.; Zheng, W.; Qian, L.; Zhang, J.; and Sun, X. 2016 · 2016
Cited alongside, same era.
Targeting Optimization for Internet Advertising by Learning from Logged Bandit Feedback
Gasparini, M.; Nuara, A.; Trovò, F.; Gatti, N.; and Restelli, M. 2018 · 2018
Later among the works it cites.
The proposal to lower P value thresholds to. 005
Ioannidis, J. P. 2018 · 2018
Later among the works it cites.
Real-time bidding with multi-agent reinforcement learning in display advertising
Jin, J.; Song, C.; Li, H.; Gai, K.; Wang, J.; and Zhang, W. 2018 · 2018
Later among the works it cites.
Data-efficient hierarchical reinforcement learning
Nachum, O.; Gu, S. S.; Lee, H.; and Levine, S. 2018 · 2018
Later among the works it cites.
A Combinatorial-Bandit Algorithm for the Online Joint Bid/Budget Optimization of Pay-per-Click Advertising Campaigns
Nuara, A.; Trovo, F.; Gatti, N.; and Restelli, M. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Kulkarni, T. D.; Narasimhan, K.; Saeedi, A.; and Tenenbaum, J. 2016 · 2016
Cited alongside, same era.
Dynamic Contextual Multi Arm Bandits in Display Advertisement
Yang, H.; and Lu, Q. 2016 · 2016
Cited alongside, same era.
Real-time bidding by reinforcement learning in display advertising
Cai, H.; Ren, K.; Zhang, W.; Malialis, K.; Wang, J.; Yu, Y.; and Guo, D. 2017 · 2017
Cited alongside, same era.
DeepFM: a factorization-machine based neural network for CTR prediction
Guo, H.; Tang, R.; Ye, Y.; Li, Z.; and He, X. 2017 · 2017
Cited alongside, same era.
Customer acquisition via display advertising using multi-armed bandit experiments
Schwartz, E. M.; Bradlow, E. T.; and Fader, P. S. 2017 · 2017
Cited alongside, same era.
Deep Reinforcement Learning for List-wise Recommendations
Zhao, X.; Zhang, L.; Ding, Z.; Yin, D.; Zhao, Y.; and Tang, J. 2017 · 2017
Cited alongside, same era.
Reinforcement Learning based Recommender System using Biclustering Technique
Choi, S.; Ha, H.; Hwang, U.; Kim, C.; Ha, J.-W.; and Yoon, S. 2018 · 2018
Cited alongside, same era.
Rohde, D.; Bonner, S.; Dunlop, T.; Vasile, F.; and Karatzoglou, A. 2018 · 2018
Later among the works it cites.
A Reinforcement Learning Framework for Explainable Recommendation
Wang, X.; Chen, Y.; Yang, J.; Wu, L.; Wu, Z.; and Xie, X. 2018b · 2018
Later among the works it cites.
DRN: A Deep Reinforcement Learning Framework for News Recommendation
Zheng, G.; Zhang, F.; Zheng, Z.; Xiang, Y.; Yuan, N. J.; Xie, X.; and Li, Z. 2018 · 2018
Later among the works it cites.
Automated Embedding Size Search in Deep Recommender Systems
Liu, H.; Zhao, X.; Wang, C.; Liu, X.; and Tang, J. 2020 · 2020
Closest in time.
Deep Reinforcement Learning for Information Retrieval: Fundamentals and Advances
Zhang, W.; Zhao, X.; Zhao, L.; Yin, D.; Yang, G. H.; and Beutel, A. 2020 · 2020
Closest in time.
Neural Interactive Collaborative Filtering
Zou, L.; Xia, L.; Gu, Y.; Zhao, X.; Liu, W.; Huang, J. X.; and Yin, D. 2020 · 2020
Closest in time.
Towards Long-term Fairness in Recommendation
Ge, Y.; Liu, S.; Gao, R.; Xian, Y.; Li, Y.; Zhao, X.; Pei, C.; Sun, F.; Ge, J.; Ou, W.; et al. 2021 · 2021
Closest in time.