Fetching the paper…
Reading the bibliography…
Reward-Weighted Regression (RWR) belongs to a family of widely known iterative Reinforcement Learning algorithms based on the Expectation-Maximization framework.
Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
Peng, X. B.; Kumar, A.; Zhang, G.; and Levine, S. 2019 · 1910
Earlier work this paper cites.
Sulle successioni di funzioni ortogonali [On Sequences of Orthogonal Functions]
Severini, C. 1910 · 1910
Earlier work this paper cites.
Conditional Markov processes
Stratonovich, R. 1960 · 1960
Earlier work this paper cites.
Self-organization of orientation sensitive cells in the striate cortex
Von der Malsburg, C. 1973 · 1973
Earlier work this paper cites.
Principles of Mathematical Analysis
Rudin, W. 1976 · 1976
Earlier work this paper cites.
Maximum likelihood from incomplete data via the EM algorithm
Dempster, A. P.; Laird, N. M.; and Rubin, D. B. 1977 · 1977
Earlier work this paper cites.
On the Convergence Properties of the EM Algorithm
Wu, C. J. 1983 · 1983
Earlier work this paper cites.
Using Expectation-Maximization for Reinforcement Learning
Dayan, P.; and Hinton, G. E. 1997 · 1997
Earlier work this paper cites.
Policy Gradient Methods for Reinforcement Learning with Function Approximation
Sutton, R. S.; McAllester, D.; Singh, S.; and Mansour, Y. 1999 · 1999
Earlier work this paper cites.
Between MDPs and Semi-MDPs: A Framework for Temporal Abstraction in Reinforcement Learning
Sutton, R. S.; Precup, D.; and Singh, S. P. 1999 · 1999
Earlier work this paper cites.
Topology
Munkres, J. 2000 · 2000
Earlier work this paper cites.
A User’s Guide to Measure Theoretic Probability
Pollard, D. 2001 · 2001
Earlier work this paper cites.
Tree-Based Batch Mode Reinforcement Learning
Ernst, D.; Geurts, P.; and Wehenkel, L. 2005 · 2005
Earlier work this paper cites.
Neural Fitted Q Iteration - First Experiences with a Data Efficient Neural Reinforcement Learning Method
Riedmiller, M. A. 2005 · 2005
Cited alongside, same era.
Fitted Q-iteration in continuous action-space MDPs
Antos, A.; Munos, R.; and Szepesvári, C. 2007 · 2007
Cited alongside, same era.
Reinforcement Learning by Reward-Weighted Regression for Operational Space Control
Peters, J.; and Schaal, S. 2007 · 2007
Cited alongside, same era.
Stochastic Approximation: A Dynamical Systems Viewpoint , volume 48
Borkar, V. S. 2008 · 2008
Cited alongside, same era.
Fitted Q-iteration by Advantage Weighted Regression
Neumann, G.; and Peters, J. 2008 · 2008
Cited alongside, same era.
Learning to Control in Operational Space
Peters, J.; and Schaal, S. 2008 · 2008
Cited alongside, same era.
Reward-Weighted Regression with Sample Reuse for Direct Policy Search in Reinforcement Learning
Hachiya, H.; Peters, J.; and Sugiyama, M. 2011 · 2011
Later among the works it cites.
Policy search for motor primitives in robotics
Kober, J.; and Peters, J. 2011 · 2011
Later among the works it cites.
Variational Inference for Policy Search in changing situations
Neumann, G. 2011 · 2011
Later among the works it cites.
Weighted Likelihood Policy Search with Model Selection
Ueno, T.; Hayashi, K.; Washio, T.; and Kawahara, Y. 2012 · 2012
Later among the works it cites.
Convergence of Probability Measures
Billingsley, P. 2013 · 2013
Later among the works it cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L. 2014 · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Episodic Reinforcement Learning by Logistic Reward-Weighted Regression
Wierstra, D.; Schaul, T.; Peters, J.; and Schmidhuber, J. 2008a · 2008
Cited alongside, same era.
Fitness Expectation Maximization
Wierstra, D.; Schaul, T.; Peters, J.; and Schmidhuber, J. 2008b · 2008
Cited alongside, same era.
Efficient Sample Reuse in EM-Based Policy Search
Hachiya, H.; Peters, J.; and Sugiyama, M. 2009 · 2009
Cited alongside, same era.
Fast Gradient-Descent Methods for Temporal-Difference Learning with Linear Function Approximation
Sutton, R. S.; Maei, H. R.; Precup, D.; Bhatnagar, S.; Silver, D.; Szepesvári, C.; and Wiewiora, E. 2009 · 2009
Cited alongside, same era.
A convergent o ( n ) o(n) temporal-difference algorithm for off-policy learning with linear function approximation
Sutton, R. S.; Maei, H. R.; and Szepesvári, C. 2009 · 2009
Cited alongside, same era.
Relative Entropy Policy Search
Peters, J.; Mülling, K.; and Altun, Y. 2010 · 2010
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
Silver, D.; Huang, A.; Maddison, C. J.; Guez, A.; Sifre, L.; van den Driessche, G.; Schrittwieser, J.; Antonoglou, I.; Panneershelvam, V.; Lanctot, M.; Dieleman, S.; Grewe, D.; Nham, J.; Kalchbrenner, N.; Sutskever, I.; Lillicrap, T. P.; Leach, M.; Kavukcuoglu, K.; Graepel, T.; and Hassabis, D. 2016 · 2016
Later among the works it cites.
Maximum a Posteriori Policy Optimisation
Abdolmaleki, A.; Springenberg, J. T.; Tassa, Y.; Munos, R.; Heess, N.; and Riedmiller, M. A. 2018b · 2018
Later among the works it cites.
Hierarchical Policy Search via Return-Weighted Density Estimation
Osa, T.; and Sugiyama, M. 2018 · 2018
Later among the works it cites.
Multi-Goal Reinforcement Learning: Challenging Robotics Environments and Request for Research
Plappert, M.; Andrychowicz, M.; Ray, A.; McGrew, B.; Baker, B.; Powell, G.; Schneider, J.; Tobin, J.; Chociej, M.; Welinder, P.; Kumar, V.; and Zaremba, W. 2018 · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. S.; and Barto, A. G. 2018 · 2018
Later among the works it cites.
Autonomous navigation of stratospheric balloons using reinforcement learning
Bellemare, M. G.; Candido, S.; Castro, P. S.; Gong, J.; Machado, M. C.; Moitra, S.; Ponda, S. S.; and Wang, Z. 2020 · 2020
Later among the works it cites.