Fetching the paper…
Reading the bibliography…
We propose RecSim, a configurable platform for authoring simulation environments for recommender systems (RSs) that naturally supports sequential interaction with users.
Memory approaches to reinforcement learning in non-Markovian domains
Long-Ji Lin and Tom. M. Mitchell · 1992
Earlier work this paper cites.
Q-learning
Christopher J. C. H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
GroupLens: Applying collaborative filtering to Usenet news
Joseph A. Konstan, Bradley N. Miller, David Maltz, Jonathan L. Herlocker, Lee R. Gordon, and John Riedl · 1997
Earlier work this paper cites.
Empirical analysis of predictive algorithms for collaborative filtering
Jack S. Breese, David Heckerman, and Carl Kadie · 1998
Earlier work this paper cites.
Stated Choice Methods: Analysis and Application
Jordan J. Louviere, David A. Hensher, and Joffre D. Swait · 2000
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Optimizing search engines using clickthrough data
Thorsten Joachims · 2002
Earlier work this paper cites.
Neural fitted Q-iteration—first experiences with a data efficient neural reinforcement learning method
Martin Riedmiller · 2005
Earlier work this paper cites.
An MDP-based recommender system
Guy Shani, David Heckerman, and Ronen I. Brafman · 2005
Earlier work this paper cites.
Probabilistic matrix factorization
Ruslan Salakhutdinov and Andriy Mnih · 2007
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E. Schapire · 2010
Earlier work this paper cites.
An empirical evaluation of Thompson sampling
Olivier Chapelle and Lihong Li · 2011
Earlier work this paper cites.
The KL-UCB algorithm for bounded stochastic bandits and beyond
Aurelien Garivier and Olivier Cappe · 2011
Earlier work this paper cites.
Critiquing-based recommenders: survey and emerging trends
Li Chen and Pearl Pu · 2012
Earlier work this paper cites.
Multiple objective optimization in recommender systems
M. Rodriguez, C. Posse, and E. Zhang · 2012
Earlier work this paper cites.
Further optimal regret bounds for Thompson sampling
Shipra Agrawal and Navin Goyal · 2013
Earlier work this paper cites.
The Arcade Learning Environment: An evaluation platform for general agents
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Modeling delayed feedback in display advertising
Olivier Chapelle · 2014
Earlier work this paper cites.
Contextual markov decision processes
Assaf Hallak, Dotan Di Castro, and Shie Mannor · 2015
Earlier work this paper cites.
Focusing on the long-term: It’s good for users and business
Henning Hohnhold, Deirdre O’Brien, and Diane Tang · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Peter Sunehag, Richard Evans, Gabriel Dulac-Arnold, Yori Zwols, Daniel Visentin, and Ben Coppin · 2015
Cited alongside, same era.
Oriol Vinyals and Quoc V. Le · 2015
Cited alongside, same era.
TensorFlow: A system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Towards conversational recommender systems
Soft actor-critic algorithms and applications
Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Data center cooling using model-predictive control
Nevena Lazic, Craig Boutilier, Tyler Lu, Eehern Wong, Binz Roy, MK Ryu, and Greg Imwalle · 2018
Later among the works it cites.
Deep Dyna-Q: Integrating planning for task-completion dialogue policy learning
Baolin Peng, Xiujun Li, Jianfeng Gao, Jingjing Liu, and Kam-Fai Wong · 2018
Later among the works it cites.
David Rohde, Stephen Bonner, Travis Dunlop, Flavian Vasile, and Alexandros Karatzoglou · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Konstantina Christakopoulou, Filip Radlinski, and Katja Hofmann · 2016
Cited alongside, same era.
The MovieLens datasets: History and context
F. Maxwell Harper and Joseph A. Konstan · 2016
Cited alongside, same era.
Fusing similarity models with Markov chains for sparse sequential recommendation
Ruining He and Julian McAuley · 2016
Cited alongside, same era.
Session-based recommendations with recurrent neural networks
Balázs Hidasi, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos Tikk · 2016
Cited alongside, same era.
Collaborative filtering bandits
Shuai Li, Alexandros Karatzoglou, and Claudio Gentile · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Discrete sequential prediction of continuous actions for deep RL
Luke Metz, Julian Ibarz, Navdeep Jaitly, and James Davidson · 2017
Cited alongside, same era.
ELF: an extensive, lightweight and flexible research platform for real-time strategy games
Yuandong Tian, Qucheng Gong, Wenling Shang, Yuxin Wu, and C. Lawrence Zitnick · 2017
Cited alongside, same era.
Yueming Sun and Yi Zhang · 2018
Later among the works it cites.
AirDialogue: An environment for goal-oriented dialogue research
Wei Wei, Quoc Le, Andrew Dai, and Jia Li · 2018
Later among the works it cites.
Practical diversified recommendations on YouTube with determinantal point processes
Mark Wilhelm, Ajith Ramanathan, Alexander Bonomo, Sagar Jain, Ed H. Chi, and Jennifer Gillenwater · 2018
Later among the works it cites.
Natural environment benchmarks for reinforcement learning
Amy Zhang, Yuxin Wu, and Joelle Pineau · 2018
Later among the works it cites.
Deep reinforcement learning for page-wise recommendations
Xiangyu Zhao, Long Xia, Liang Zhang, Zhuoye Ding, Dawei Yin, and Jiliang Tang · 2018
Later among the works it cites.
DRN: A deep reinforcement learning framework for news recommendation
Guanjie Zheng, Fuzheng Zhang, Zihan Zheng, Yang Xiang, Nicholas Jing Yuan, Xing Xie, and Zhenhui Li · 2018
Later among the works it cites.
Reinforcement learning when all actions are not always available
Yash Chandak, Georgios Theocharous, Blossom Metevier, and Philip S. Thomas · 2019
Closest in time.
Challenges of real-world reinforcement learning
Gabriel Dulac-Arnold, Daniel Mankowitz, and Todd Hester · 2019
Closest in time.
TF-Agents: A library for reinforcement learning in tensorflow
Sergio Guadarrama, Anoop Korattikara, Oscar Ramirez, Pablo Castro, Ethan Holly, Sam Fishman, Ke Wang, Ekaterina Gonina, Neal Wu, Chris Harris, Vincent Vanhoucke, and Eugene Brevdo · 2019
Closest in time.
SlateQ: A tractable decomposition for reinforcement learning with recommendation sets
Eugene Ie, Vihan Jain, Jing Wang, Sanmit Narvekar, Ritesh Agarwal, Rui Wu, Heng-Tze Cheng, Tushar Chandra, and Craig Boutilier · 2019
Closest in time.
Beyond greedy ranking: Slate optimization via List-CVAE
Ray Jiang, Sven Gowal, Timothy A. Mann, and Danilo J. Rezende · 2019
Closest in time.
Building personalized simulator for interactive search
Qianlong Liu, Baoliang Cui, Zhongyu Wei, Baolin Peng, Haikuan Huang, Hongbo Deng, Jianye Hao, Xuanjing Huang, and Kam-Fai Wong · 2019
Closest in time.
Advantage amplification in slowly evolving latent-state environments
Martin Mladenov, Ofer Meshi, Jayden Ooi, Dale Schuurmans, and Craig Boutilier · 2019
Closest in time.
Virtual-taobao: Virtualizing real-world online retail environment for reinforcement learning
Jing-Cheng Shi, Yang Yu, Qing Da, Shi-Yong Chen, and An-Xiang Zeng · 2019
Closest in time.
Toward simulating environments in reinforcement learning based recommendations
Xiangyu Zhao, Long Xia, Zhuoye Ding, Dawei Yin, and Jiliang Tang · 2019
Closest in time.
Agenda-based user simulation for bootstrapping a POMDP dialogue system
Jost Schatzmann, Blaise Thomson, Karl Weilhammer, Hui Ye, and Steve Young · 2038
Closest in time.