Fetching the paper…
Reading the bibliography…
Multi-task reinforcement learning (RL) aims to simultaneously learn policies for solving many tasks.
Learning to achieve goals
Kaelbling, L. P · 1993
Earlier work this paper cites.
Multitask learning
Caruana, R · 1997
Earlier work this paper cites.
Using expectation-maximization for reinforcement learning
Dayan, P. and Hinton, G. E · 1997
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y · 2004
Earlier work this paper cites.
Elements of information theory (wiley series in telecommunications and signal processing), 2006
Cover, T. M. and Thomas, J. A · 2006
Earlier work this paper cites.
Maximum margin planning
Ratliff, N. D., Bagnell, J. A., and Zinkevich, M. A · 2006
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
Peters, J. and Schaal, S · 2007
Earlier work this paper cites.
Linearly-solvable markov decision problems
Todorov, E · 2007
Earlier work this paper cites.
General duality between optimal control and estimation
Todorov, E · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A., Bagnell, J. A., and Dey, A. K · 2008
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
Toussaint, M · 2009
Earlier work this paper cites.
Optimal control as a graphical model inference problem
Kappen, H. J., Gómez, V., and Opper, M · 2012
Earlier work this paper cites.
Relative entropy and free energy dualities: Connections to path integral and kl control
Theodorou, E. A. and Todorov, E · 2012
Earlier work this paper cites.
Learning to learn
Thrun, S. and Pratt, L · 2012
Earlier work this paper cites.
Variational policy search via trajectory optimization
Levine, S. and Koltun, V · 2013
Earlier work this paper cites.
On stochastic optimal control and reinforcement learning by approximate inference
Rawlik, K., Toussaint, M., and Vijayakumar, S · 2013
Cited alongside, same era.
Shared autonomy via hindsight optimization
Javdani, S., Srinivasa, S. S., and Bagnell, J. A · 2015
Cited alongside, same era.
Actor-mimic: Deep multitask and transfer reinforcement learning
Parisotto, E., Ba, J. L., and Salakhutdinov, R · 2015
Cited alongside, same era.
Rusu, A. A., Colmenarejo, S. G., Gulcehre, C., Desjardins, G., Kirkpatrick, J., Pascanu, R., Mnih, V., Kavukcuoglu, K., and Hadsell, R · 2015
Cited alongside, same era.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Cited alongside, same era.
Tf-agents: A library for reinforcement learning in tensorflow, 2018
Guadarrama, S., Korattikara, A., Ramirez, O., Castro, P., Holly, E., Fishman, S., Wang, K., Gonina, E., Harris, C., Vanhoucke, V., et al · 2018
Later among the works it cites.
Recurrent experience replay in distributed reinforcement learning
Kapturowski, S., Ostrovski, G., Quan, J., Munos, R., and Dabney, W · 2018
Later among the works it cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Levine, S · 2018
Later among the works it cites.
Oh, J., Guo, Y., Singh, S., and Lee, H · 2018
Later among the works it cites.
Temporal difference models: Model-free deep rl for model-based control
Pong, V., Gu, S., Dalal, M., and Levine, S · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Guided cost learning: Deep inverse optimal control via policy optimization
Finn, C., Levine, S., and Abbeel, P · 2016
Cited alongside, same era.
Neural symbolic machines: Learning semantic parsers on freebase with weak supervision
Liang, C., Berant, J., Le, Q., Forbus, K. D., and Lao, N · 2016
Cited alongside, same era.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, O. P., and Zaremba, W · 2017
Cited alongside, same era.
Divide-and-conquer reinforcement learning
Ghosh, D., Singh, A., Rajeswaran, A., Kumar, V., and Levine, S · 2017
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Inverse reward design
Hadfield-Menell, D., Milli, S., Abbeel, P., Russell, S. J., and Dragan, A · 2017
Cited alongside, same era.
Distral: Robust multitask reinforcement learning
Teh, Y., Bapst, V., Czarnecki, W. M., Quan, J., Kirkpatrick, J., Hadsell, R., Heess, N., and Pascanu, R · 2017
Cited alongside, same era.
Later among the works it cites.
Learning by playing-solving sparse reward tasks from scratch
Riedmiller, M., Hafner, R., Lampe, T., Neunert, M., Degrave, J., Van de Wiele, T., Mnih, V., Heess, N., and Springenberg, J. T · 2018
Later among the works it cites.
Semi-parametric topological memory for navigation
Savinov, N., Dosovitskiy, A., and Koltun, V · 2018
Later among the works it cites.
Search on the replay buffer: Bridging planning and reinforcement learning
Eysenbach, B., Salakhutdinov, R., and Levine, S · 2019
Later among the works it cites.
Learning to reach goals without reinforcement learning
Ghosh, D., Gupta, A., Fu, J., Reddy, A., Devine, C., Eysenbach, B., and Levine, S · 2019
Later among the works it cites.
Multi-task deep reinforcement learning with popart
Hessel, M., Soyer, H., Espeholt, L., Czarnecki, W., Schmitt, S., and van Hasselt, H · 2019
Later among the works it cites.
Learning latent plans from play
Lynch, C., Khansari, M., Xiao, T., Kumar, V., Tompson, J., Levine, S., and Sermanet, P · 2019
Later among the works it cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Peng, X. B., Kumar, A., Zhang, G., and Levine, S · 2019
Later among the works it cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Rakelly, K., Zhou, A., Quillen, D., Finn, C., and Levine, S · 2019
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2019
Later among the works it cites.
Gradient surgery for multi-task learning
Yu, T., Kumar, S., Gupta, A., Levine, S., Hausman, K., and Finn, C · 2020
Closest in time.