Fetching the paper…
Reading the bibliography…
Most model-free reinforcement learning methods leverage state representations (embeddings) for generalization, but either ignore structure in the space of actions or assume the structure is provided a priori.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximationtechnical
Tsitsiklis, J. and Van Roy, B · 1996
Earlier work this paper cites.
The actor-critic algorithm as multi-time-scale stochastic approximation
Borkar, V. S. and Konda, V. R · 1997
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N · 2000
Earlier work this paper cites.
Motor learning through the combination of primitives
Mussa-Ivaldi, F. A. and Bizzi, E · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
An introduction to hidden markov models and bayesian networks
Ghahramani, Z · 2001
Earlier work this paper cites.
Learning attractor landscapes for learning motor primitives
Ijspeert, A. J., Nakanishi, J., and Schaal, S · 2003
Earlier work this paper cites.
The construction of movement with behavior-specific and behavior-independent modules
Jing, J., Cropper, E. C., Hurwitz, I., and Weiss, K. R · 2004
Earlier work this paper cites.
Modularity of motor output evoked by intraspinal microstimulation in cats
Lemay, M. A. and Grill, W. M · 2004
Earlier work this paper cites.
Reinforcement learning with factored states and actions
Sallans, B. and Hinton, G. E · 2004
Earlier work this paper cites.
An MDP-based recommender system
Shani, G., Heckerman, D., and Brafman, R. I · 2005
Earlier work this paper cites.
Integrating affect sensors in an intelligent tutoring system
Sidney, K. D., Craig, S. D., Gholson, B., Franklin, S., Picard, R., and Graesser, A. C · 2005
Earlier work this paper cites.
Dynamic movement primitives-a framework for motor control in humans and humanoid robotics
Schaal, S · 2006
Earlier work this paper cites.
Natural actor–critic algorithms
Bhatnagar, S., Sutton, R. S., Ghavamzadeh, M., and Lee, M · 2009
Earlier work this paper cites.
Stochastic approximation: a dynamical systems viewpoint , volume 48
Borkar, V. S · 2009
Earlier work this paper cites.
Learning motor primitives for robotics
Kober, J. and Peters, J · 2009
Cited alongside, same era.
Using continuous action spaces to solve discrete problems
Van Hasselt, H. and Wiering, M. A · 2009
Cited alongside, same era.
Value function approximation in reinforcement learning using the fourier basis
Konidaris, G., Osentoski, S., and Thomas, P. S · 2011
Cited alongside, same era.
Generalized value functions for large action sets
Pazis, J. and Parr, R · 2011
Cited alongside, same era.
Policy gradient coagent networks
Thomas, P. S · 2011
Cited alongside, same era.
Conjugate Markov decision processes
Thomas, P. S. and Barto, A. G · 2011
Cited alongside, same era.
Online symbolic gradient-based optimization for factored action mdps
Cui, H. and Khardon, R · 2016
Later among the works it cites.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K · 2016
Later among the works it cites.
Deep learning without poor local minima
Kawaguchi, K · 2016
Later among the works it cites.
Loss is its own reward: Self-supervision for reinforcement learning
Shelhamer, E., Mahmoudieh, P., Argus, M., and Darrell, T · 2016
Later among the works it cites.
Reinforcement learning for electric power system decision and control: Past considerations and perspectives
Glavic, M., Fonteneau, R., and Ernst, D · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Degris, T., White, M., and Sutton, R. S · 2012
Cited alongside, same era.
Motor primitive discovery
Thomas, P. S. and Barto, A. G · 2012
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J · 2013
Cited alongside, same era.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M · 2014
Cited alongside, same era.
Bias in natural actor-critic algorithms
Thomas, P · 2014
Cited alongside, same era.
Global optimality in neural network training
Haeffele, B. D. and Vidal, R · 2017
Later among the works it cites.
A deep reinforcement learning framework for the financial portfolio management problem
Jiang, Z., Xu, D., and Liang, J · 2017
Later among the works it cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
Sharma, S., Suresh, A., Ramesh, R., and Ravindran, B · 2017
Later among the works it cites.
Lifted stochastic planning, belief propagation and marginal MAP
Cui, H. and Khardon, R · 2018
Later among the works it cites.
Combined reinforcement learning via abstract representations
François-Lavet, V., Bengio, Y., Precup, D., and Pineau, J · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
The global optimization geometry of shallow linear neural networks
Zhu, Z., Soudry, D., Eldar, Y. C., and Wakin, M. B · 2018
Later among the works it cites.
The natural language of actions
Tennenholtz, G. and Mannor, S · 2019
Closest in time.