Fetching the paper…
Reading the bibliography…
We present a novel podcast recommender system deployed at industrial scale.
Eugene Ie, Vihan Jain, Jing Wang, Sanmit Narvekar, Ritesh Agarwal, Rui Wu, Heng-Tze Cheng, Morgane Lustman, Vince Gatto, Paul Covington, et al · 1905
Earlier work this paper cites.
Recsim: A configurable simulation platform for recommender systems
Eugene Ie, Chih-wei Hsu, Martin Mladenov, Vihan Jain, Sanmit Narvekar, Jing Wang, Rui Wu, and Craig Boutilier · 1909
Earlier work this paper cites.
Estimating causal effects of treatments in randomized and nonrandomized studies
Donald B Rubin · 1974
Earlier work this paper cites.
Approximations of dynamic programs, I
Ward Whitt · 1978
Earlier work this paper cites.
Aggregation in dynamic programming
James C Bean, John R Birge, and Robert L Smith · 1987
Earlier work this paper cites.
Surrogate endpoints in clinical trials: definition and operational criteria
Ross L Prentice · 1989
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Stable function approximation in dynamic programming
Geoffrey J Gordon · 1995
Earlier work this paper cites.
Reinforcement learning with soft state aggregation
Satinder P Singh, Tommi Jaakkola, and Michael I Jordan · 1995
Earlier work this paper cites.
Feature-based methods for large scale dynamic programming
John N Tsitsiklis and Benjamin Van Roy · 1996
Earlier work this paper cites.
Using randomization to break the curse of dimensionality
John Rust · 1997
Earlier work this paper cites.
Optimal mailing of catalogs: a new methodology using estimable structural dynamic programming models
Füsun Gönül and Meng Ze Shi · 1998
Earlier work this paper cites.
Actor-critic algorithms
Vijay Konda and John Tsitsiklis · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Internet recommendation systems
Asim Ansari, Skander Essegaier, and Rajeev Kohli · 2000
Earlier work this paper cites.
Content-based book recommending using learning for text categorization
Raymond J Mooney and Loriene Roy · 2000
Earlier work this paper cites.
Modeling customer relationships as markov chains
Phillip E Pfeifer and Robert L Carraway · 2000
Earlier work this paper cites.
Comments on the origin and application of markov decision processes
Ronald A Howard · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
An mdp-based recommender system
Guy Shani, David Heckerman, Ronen I Brafman, and Craig Boutilier · 2005
Earlier work this paper cites.
Improving recommendation lists through topic diversification
Cai-Nicolas Ziegler, Sean M McNee, Joseph A Konstan, and Georg Lausen · 2005
Earlier work this paper cites.
How to compute optimal catalog mailing decisions
Füsun F Gönül and Frenkel Ter Hofstede · 2006
Earlier work this paper cites.
Towards a unified theory of state abstraction for MDPs
Lihong Li, Thomas J Walsh, and Michael L Littman · 2006
Earlier work this paper cites.
Marketing models of service and relationships
Roland T Rust and Tuck Siong Chung · 2006
Earlier work this paper cites.
Dynamic catalog mailing policies
Duncan I Simester, Peng Sun, and John N Tsitsiklis · 2006
Cited alongside, same era.
A hidden markov model of customer relationship dynamics
Oded Netzer, James M Lattin, and Vikram Srinivasan · 2008
Cited alongside, same era.
Matrix factorization techniques for recommender systems
Yehuda Koren, Robert Bell, and Chris Volinsky · 2009
Cited alongside, same era.
Causal inference in statistics: An overview
Judea Pearl · 2009
Cited alongside, same era.
A survey of collaborative filtering techniques
Xiaoyuan Su and Taghi M Khoshgoftaar · 2009
Cited alongside, same era.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E Schapire · 2010
Cited alongside, same era.
The surrogate index: Combining short-term proxies to estimate long-term treatment effects more rapidly and precisely
Susan Athey, Raj Chetty, Guido W Imbens, and Hyunseung Kang · 2019
Later among the works it cites.
Environment reconstruction with hidden confounders for reinforcement learning based recommendation
Wenjie Shang, Yang Yu, Qingyang Li, Zhiwei Qin, Yiping Meng, and Jieping Ye · 2019
Later among the works it cites.
Reinforcement learning to optimize long-term user engagement in recommender systems
Lixin Zou, Long Xia, Zhuoye Ding, Jiaxing Song, Weidong Liu, and Dawei Yin · 2019
Later among the works it cites.
Algorithmic effects on the diversity of consumption on Spotify
Ashton Anderson, Lucas Maystre, Ian Anderson, Rishabh Mehrotra, and Mounia Lalmas · 2020
Later among the works it cites.
Fatigue-aware bandits for dependent click models
Junyu Cao, Wei Sun, Zuo-Jun Max Shen, and Markus Ettl · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dynamic allocation of pharmaceutical detailing and sampling for long-term profitability
Ricardo Montoya, Oded Netzer, and Kamel Jedidi · 2010
Cited alongside, same era.
Solving the apparent diversity-accuracy dilemma of recommender systems
Tao Zhou, Zoltán Kuscsik, Jian-Guo Liu, Matúš Medo, Joseph Rushton Wakeling, and Yi-Cheng Zhang · 2010
Cited alongside, same era.
Media exposure through the funnel: A model of multi-stage attribution
Vibhanshu Abhishek, Peter Fader, and Kartik Hosanagar · 2012
Cited alongside, same era.
Dynamic programming and optimal control: Volume I , volume 1
Dimitri Bertsekas · 2012
Cited alongside, same era.
A joint model of usage and churn in contractual settings
Eva Ascarza and Bruce GS Hardie · 2013
Cited alongside, same era.
Abstraction selection in model-based reinforcement learning
Nan Jiang, Alex Kulesza, and Satinder Singh · 2015
Cited alongside, same era.
Daniel Russo · 2020
Later among the works it cites.
Self-supervised reinforcement learning for recommender systems
Xin Xin, Alexandros Karatzoglou, Ioannis Arapakis, and Joemon M Jose · 2020
Later among the works it cites.
Targeting for long-term outcomes
Jeremy Yang, Dean Eckles, Paramveer Dhillon, and Sinan Aral · 2020
Later among the works it cites.
Reinforcement learning based recommender systems: A survey
M Mehdi Afsar, Trafford Crump, and Behrouz Far · 2021
Later among the works it cites.
Challenges of real-world reinforcement learning: definitions, benchmarks and analysis
Gabriel Dulac-Arnold, Nir Levine, Daniel J Mankowitz, Jerry Li, Cosmin Paduraru, Sven Gowal, and Todd Hester · 2021
Later among the works it cites.
Reinforcement learning, bit by bit
Xiuyuan Lu, Benjamin Van Roy, Vikranth Dwaracherla, Morteza Ibrahimi, Ian Osband, and Zheng Wen · 2021
Later among the works it cites.
Deep learning for recommender systems: A netflix case study
Harald Steck, Linas Baltrunas, Ehtsham Elahi, Dawen Liang, Yves Raimond, and Justin Basilico · 2021
Later among the works it cites.
A survey on session-based recommender systems
Shoujin Wang, Longbing Cao, Yan Wang, Quan Z Sheng, Mehmet A Orgun, and Defu Lian · 2021
Later among the works it cites.
Hierarchical reinforcement learning for integrated recommendation
Ruobing Xie, Shaoliang Zhang, Rui Wang, Feng Xia, and Leyu Lin · 2021
Later among the works it cites.
Learning personalized product recommendations with customer disengagement
Hamsa Bastani, Pavithra Harsha, Georgia Perakis, and Divya Singhvi · 2022
Later among the works it cites.
Modeling attrition in recommender systems with departing bandits
Omer Ben-Porat, Lee Cohen, Liu Leqi, Zachary C Lipton, and Yishay Mansour · 2022
Later among the works it cites.
Off-policy actor-critic for recommender systems
Minmin Chen, Can Xu, Vince Gatto, Devanshu Jain, Aviral Kumar, and Ed Chi · 2022
Later among the works it cites.
Reinforcement learning for budget constrained recommendations
Ehtsham Elahi, James McInerney, Nathan Kallus, Dario Garcia Garcia, and Justin Basilico · 2022
Later among the works it cites.
Morphing for consumer dynamics: Bandits meet hidden markov models
Gui Liberali and Alina Ferecatu · 2022
Later among the works it cites.
Shapley meets uniform: An axiomatic framework for attribution in online advertising
Raghav Singal, Omar Besbes, Antoine Desir, Vineet Goyal, and Garud Iyengar · 2022
Later among the works it cites.
Surrogate for long-term user experience in recommender systems
Yuyan Wang, Mohit Sharma, Can Xu, Sriraj Badam, Qian Sun, Lee Richardson, Lisa Chung, Ed H Chi, and Minmin Chen · 2022
Later among the works it cites.
Multi-granularity fatigue in recommendation
Ruobing Xie, Cheng Ling, Shaoliang Zhang, Feng Xia, and Leyu Lin · 2022
Later among the works it cites.
On the statistical benefits of temporal difference learning
David Cheikhi and Daniel Russo · 2023
Closest in time.