Fetching the paper…
Reading the bibliography…
Prior work has proposed a simple strategy for reinforcement learning (RL): label experience with the outcomes achieved in that experience, and then imitate the relabeled experience.
Kumar, A., Peng, X. B., and Levine, S. (2019) · 1912
Earlier work this paper cites.
Training agents using upside-down reinforcement learning
Srivastava, R. K., Shyam, P., Mutz, F., Jaśkowski, W., and Schmidhuber, J. (2019) · 1912
Earlier work this paper cites.
Learning to achieve goals
Kaelbling, L. P. (1993) · 1993
Earlier work this paper cites.
Neuro-dynamic programming
Bertsekas, D. P. and Tsitsiklis, J. N. (1996) · 1996
Earlier work this paper cites.
Using expectation-maximization for reinforcement learning
Dayan, P. and Hinton, G. E. (1997) · 1997
Earlier work this paper cites.
Generalized hindsight for reinforcement learning
Li, A. C., Pinto, L., and Abbeel, P. (2020a) · 2002
Earlier work this paper cites.
Lynch, C. and Sermanet, P. (2020) · 2005
Earlier work this paper cites.
Probabilistic inference for solving (PO)MDPs
Toussaint, M., Harmeling, S., and Storkey, A. (2006) · 2006
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
Peters, J. and Schaal, S. (2007) · 2007
Earlier work this paper cites.
Fitted Q-iteration by advantage weighted regression
Neumann, G., Peters, J., et al. (2009) · 2008
Earlier work this paper cites.
Learning motor primitives for robotics
Kober, J. and Peters, J. (2009) · 2009
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
Toussaint, M. (2009) · 2009
Earlier work this paper cites.
Relative entropy policy search
Peters, J., Mulling, K., and Altun, Y. (2010) · 2010
Earlier work this paper cites.
Expectation-maximization methods for solving (po) mdps and optimal control problems
Toussaint, M., Storkey, A., and Harmeling, S. (2010) · 2010
Cited alongside, same era.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Ziebart, B. D. (2010) · 2010
Cited alongside, same era.
Variational inference for policy search in changing situations
Neumann, G. et al. (2011) · 2011
Cited alongside, same era.
Planning from pixels using inverse dynamics models
Paster, K., McIlraith, S. A., and Ba, J. (2020) · 2012
Cited alongside, same era.
Variational policy search via trajectory optimization
Levine, S. and Koltun, V. (2013) · 2013
Cited alongside, same era.
Multi-objective reinforcement learning using sets of pareto dominating policies
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
Goal-conditioned imitation learning
Ding, Y., Florensa, C., Abbeel, P., and Phielipp, M. (2019) · 2019
Later among the works it cites.
A theory of regularized markov decision processes
Geist, M., Scherrer, B., and Pietquin, O. (2019) · 2019
Later among the works it cites.
Policy continuation with hindsight inverse dynamics
Sun, H., Li, Z., Liu, X., Zhou, B., and Lin, D. (2019) · 2019
Later among the works it cites.
Rewriting history with inverse rl: Hindsight inference for policy improvement
Eysenbach, B., Geng, X., Levine, S., and Salakhutdinov, R. (2020) · 2020
Later among the works it cites.
Learning to reach goals via iterated supervised learning
Ghosh, D., Gupta, A., Reddy, A., Fu, J., Devin, C. M., Eysenbach, B., and Levine, S. (2020) · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Van Moffaert, K. and Nowé, A. (2014) · 2014
Cited alongside, same era.
Multi-objective deep reinforcement learning
Mossalam, H., Assael, Y. M., Roijers, D. M., and Whiteson, S. (2016) · 2016
Cited alongside, same era.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, O. P., and Zaremba, W. (2017) · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Cited alongside, same era.
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D. (2018) · 2018
Cited alongside, same era.
Self-imitation learning
Oh, J., Guo, Y., Singh, S., and Lee, H. (2018) · 2018
Cited alongside, same era.
Semi-parametric topological memory for navigation
Savinov, N., Dosovitskiy, A., and Koltun, V. (2018) · 2018
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I. (2021) · 2021
Later among the works it cites.
Rvs: What is essential for offline rl via supervised learning?
Emmons, S., Eysenbach, B., Kostrikov, I., and Levine, S. (2021) · 2021
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
Kostrikov, I., Nair, A., and Levine, S. (2021) · 2021
Later among the works it cites.
Language conditioned imitation learning over unstructured data
Lynch, C. and Sermanet, P. (2021) · 2021
Later among the works it cites.
Interactive learning from activity description
Nguyen, K. X., Misra, D., Schapire, R., Dudík, M., and Shafto, P. (2021) · 2021
Later among the works it cites.
MHER: Model-based hindsight experience replay
Yang, R., Fang, M., Han, L., Du, Y., Luo, F., and Li, X. (2021) · 2021
Later among the works it cites.