Fetching the paper…
Reading the bibliography…
In many practical uses of reinforcement learning (RL) the set of actions available at a given state is a random variable, with realizations governed by an exogenous stochastic process.
A polynomial time bound for Howard’s policy improvement algorithm
U. Meister and U. Holzbaur · 1986
Earlier work this paper cites.
Geometric Algorithms and Combinatorial Optimization
Martin Grötschel, Lászlo Lovász, and Alexander Schrijver · 1988
Earlier work this paper cites.
Solving h-horizon, stationary Markov decision problems in time proportional to log(h)
Paul Tseng · 1990
Earlier work this paper cites.
Theoretical Computer Science
Shortest paths without a map · 1991
Earlier work this paper cites.
Q-learning
Christopher J. C. H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
On targeting Markov segments
Moses Charikar, Ravi Kumar, Prabhakar Raghavan, Sridhar Rajagopalan, and Andrew Tomkins · 1999
Earlier work this paper cites.
Route planning under uncertainty: The Canadian traveller problem
Evdokia Nikolova and David R. Karger · 2008
Earlier work this paper cites.
Sleeping experts and bandits with stochastic action availability and adversarial rewards
Varun Kanade, H Brendan McMahan, and Brent Bryan · 2009
Cited alongside, same era.
A Markov chain model for integrating behavioral targeting into contextual advertising
Ting Li, Ning Liu, Jun Yan, Gang Wang, Fengshan Bai, and Zheng Chen · 2009
Cited alongside, same era.
Mining advertiser-specific user behavior using adfactors
Nikolay Archak, Vahab S. Mirrokni, and S. Muthukrishnan · 2010
Cited alongside, same era.
Regret bounds for sleeping experts and bandits
Robert Kleinberg, Alexandru Niculescu-Mizil, and Yogeshwer Sharma · 2010
Cited alongside, same era.
The simplex and policy-iteration methods are strongly polynomial for the Markov decision problem with a fixed discount rate
Yinyu Ye · 2011
Cited alongside, same era.
Budget optimization for sponsored search: Censored learning in MDPs
Strategy iteration is strongly polynomial for 2-player turn-based stochastic games with a constant discount factor
Thomas Dueholm Hansen, Peter Bro Miltersen, and Uri Zwick · 2013
Later among the works it cites.
Playing Atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Later among the works it cites.
Concurrent reinforcement learning from customer interactions
David Silver, Leonard Newnham, David Barker, Suzanne Weller, and Jason McFall · 2013
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Later among the works it cites.
Personalized ad recommendation systems for life-time value optimization with guarantees
Georgios Theocharous, Philip S. Thomas, and Mohammad Ghavamzadeh · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kareem Amin, Michael Kearns, Peter Key, and Anton Schwaighofer · 2012
Cited alongside, same era.
Budget optimization for online campaigns with positive carryover effects
Nikolay Archak, Vahab Mirrokni, and S. Muthukrishnan · 2012
Cited alongside, same era.
Stochastic shortest path problems with recourse
George H. Polychronopoulos and John N. Tsitsiklis
Cited in the paper.
Later among the works it cites.
Logistic Markov decision processes
Martin Mladenov, Craig Boutilier, Dale Schuurmans, Ofer Meshi, Gal Elidan, and Tyler Lu · 2017
Later among the works it cites.
Planning and learning in Markov decision processes with stochastic action sets
Craig Boutilier, Alon Cohen, Avinatan Hassidim, Yishay Mansour, Ofer Meshi, Martin Mladenov, and Dale Schuurmans · 2018
Closest in time.