Fetching the paper…
Reading the bibliography…
Being able to reason in an environment with a large number of discrete actions is essential to bringing reinforcement learning to a larger class of problems.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Lin, Long-Ji · 1992
Earlier work this paper cites.
Solving multiclass learning problems via error-correcting output codes
Dietterich, Thomas G. and Bakiri, Ghulum · 1995
Earlier work this paper cites.
Adaptive critic designs
Prokhorov, Danil V, Wunsch, Donald C, et al · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Sutton, Richard S and Barto, Andrew G · 1998
Earlier work this paper cites.
Reinforcement learning as classification: Leveraging modern classifiers
Lagoudakis, Michail and Parr, Ronald · 2003
Earlier work this paper cites.
Using continuous action spaces to solve discrete problems
Van Hasselt, Hado, Wiering, Marco, et al · 2009
Earlier work this paper cites.
Reinforcement learning in feedback control
Hafner, Roland and Riedmiller, Martin · 2011
Cited alongside, same era.
Generalized value functions for large action sets
Pazis, Jason and Parr, Ron · 2011
Cited alongside, same era.
Fast reinforcement learning with large action sets using error-correcting output codes for mdp factorization
Dulac-Arnold, Gabriel, Denoyer, Ludovic, Preux, Philippe, and Gallinari, Patrick · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, Emanuel, Erez, Tom, and Tassa, Yuval · 2012
Cited alongside, same era.
Scalable nearest neighbor algorithms for high dimensional data
Muja, Marius and Lowe, David G · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
Silver, David, Lever, Guy, Heess, Nicolas, Degris, Thomas, Wierstra, Daan, and Riedmiller, Martin · 2014
Later among the works it cites.
Deep reinforcement learning with an unbounded action space
He, Ji, Chen, Jianshu, He, Xiaodong, Gao, Jianfeng, Li, Lihong, Deng, Li, and Ostendorf, Mari · 2015
Closest in time.
Continuous control with deep reinforcement learning
Lillicrap, Timothy P, Hunt, Jonathan J, Pritzel, Alexander, Heess, Nicolas, Erez, Tom, Tassa, Yuval, Silver, David, and Wierstra, Daan · 2015
Closest in time.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, et al · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sunehag, Peter, Evans, Richard, Dulac-Arnold, Gabriel, Zwols, Yori, Visentin, Daniel, and Coppin, Ben · 2015
Closest in time.