Fetching the paper…
Reading the bibliography…
Many real-world problems come with action spaces represented as feature vectors.
Prospect theory: An analysis of decisions under risk
Daniel Kahneman and Amos Tversky · 1979
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
M. Puterman · 1994
Earlier work this paper cites.
The use of MMR, diversity-based reranking for reordering documents and producing summaries
Jaime Carbonell and Jade Goldstein · 1998
Earlier work this paper cites.
Reinforcement Learning
R. Sutton and A. Barto · 1998
Earlier work this paper cites.
Cumulated gain-based evaluation of ir techniques
K. Järvelin and J. Kekäläinen · 2002
Earlier work this paper cites.
Autonomous inverted helicopter flight via reinforcement learning
A. Ng, A. Coates, M. Diel, V. Ganapathi, J. Schulte, B. Tse, E. Berger, and E. Liang · 2004
Earlier work this paper cites.
Submodular Functions and Optimization: Second Edition
S. Fujishige · 2005
Earlier work this paper cites.
An MDP-based recommender system
G. Shani, RI. Brafman and D, Heckerman · 2005
Cited alongside, same era.
Using continuous action spaces to solve discrete problems
Hado Van Hasselt, Marco Wiering, et al · 2009
Cited alongside, same era.
Search engines: information retrieval in practice
W. Bruce Croft, Donald Metzler, and Trevor Strohman · 2010
Cited alongside, same era.
Non-Stochastic Bandit Slate Problems
S. Kale, L. Reyzin, and R. Schapire · 2010
Cited alongside, same era.
Artificial Intelligence: A Modern Approach
S. J. Russell and P. Norvig · 2010
Cited alongside, same era.
Non-deterministic policies in markovian decision processes
M.M. Fard and J. Pineau · 2011
Cited alongside, same era.
A literature review and classification of recommender systems research
Deuk Hee Park, Hyea Kyeong Kim, Il Young Choi, and Jae Kyeong Kim · 2012
Later among the works it cites.
Matroid bandits: Fast combinatorial optimization with learning
B. Kveton, Z. Wen, A. Ashkan, H. Eydgahi, and B. Eriksson · 2014
Later among the works it cites.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin A. Riedmiller · 2014
Later among the works it cites.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Closest in time.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Rusu, J. Veness, M. Bellemare, A. Graves, M. Riedmiller, A. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Yue and C. Guestrin · 2011
Cited alongside, same era.
Closest in time.
Fast reinforcement learning in large discrete action spaces
Gabriel Dulac-Arnold, Richard Evans, and Peter Sunehag · 2016
Closest in time.