Understand
In this paper reinforcement learning with binary vector actions was investigated.
- We suggest an effective architecture of the neural networks for approximating an action-value function with binary vector actions.
- The proposed architecture approximates the action-value function by a linear function with respect to the action vector, but is still non-linear with respect to the state input.
- We show that this approximation method enables the efficient calculation of greedy action selection and softmax action selection.
Built on
Information processing in dynamical systems: foundations of harmony theory
P Smolensky · 1986
Earlier work this paper cites.
Learning from delayed rewards
Christopher John Cornish Hellaby Watkins · 1989
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Long-Ji Lin · 1992
Earlier work this paper cites.
Similar
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 1998
Cited alongside, same era.
Using free energies to represent q-values in a multiagent reinforcement learning task
Brian Sallans and Geoffrey E Hinton · 2000
Cited alongside, same era.
Reinforcement learning with factored states and actions
Brian Sallans and Geoffrey E Hinton · 2004
Cited alongside, same era.
Reinforcement learning of motor skills in high dimensions: A path integral approach
Evangelos Theodorou, Jonas Buchli, and Stefan Schaal · 2010
Cited alongside, same era.
Then
Actor-critic reinforcement learning with energy-based policies
Nicolas Heess, David Silver, and Yee Whye Teh · 2012
Later among the works it cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Closest in time.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Closest in time.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…