Fetching the paper…
Reading the bibliography…
The ability to effectively reuse prior knowledge is a key requirement when building general and flexible Reinforcement Learning (RL) agents.
Finding structure in reinforcement learning
Sebastian Thrun and Anton Schwartz · 1994
Earlier work this paper cites.
Reusing learned policies between similar problems
Mike Bowling and Manuela Veloso · 1998
Earlier work this paper cites.
Reusing old policies to accelerate learning on new mdps
Daniel S Bernstein · 1999
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Policyblocks: An algorithm for creating useful macro-actions in reinforcement learning
Marc Pickett and Andrew G Barto · 2002
Earlier work this paper cites.
Learning movement primitives
Stefan Schaal, Jan Peters, Jun Nakanishi, and Auke Ijspeert · 2005
Earlier work this paper cites.
Relational macros for transfer in reinforcement learning
Lisa Torrey, Jude Shavlik, Trevor Walker, and Richard Maclin · 2007
Earlier work this paper cites.
The utility of temporal abstraction in reinforcement learning
Nicholas K Jong, Todd Hester, and Peter Stone · 2008
Earlier work this paper cites.
Importance weighted policy learning and adaption
Alexandre Galashov, Jakub Sygnowski, Guillaume Desjardins, Jan Humplik, Leonard Hasenclever, Rae Jeong, Yee Whye Teh, and Nicolas Heess · 2009
Earlier work this paper cites.
An efficient orientation filter for inertial and inertial/magnetic sensor arrays
Sebastian Madgwick et al · 2010
Earlier work this paper cites.
Clustering via dirichlet process mixture models for portable skill discovery
Scott Niekum and Andrew G Barto · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Hierarchical relative entropy policy search
Christian Daniel, Gerhard Neumann, and Jan Peters · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Dynamical movement primitives: learning attractor models for motor behaviors
Auke Jan Ijspeert, Jun Nakanishi, Heiko Hoffmann, Peter Pastor, and Stefan Schaal · 2013
Earlier work this paper cites.
Learning to select and generalize striking movements in robot table tennis
Katharina Mülling, Jens Kober, Oliver Kroemer, and Jan Peters · 2013
Earlier work this paper cites.
Probabilistic movement primitives
Alexandros Paraschos, Christian Daniel, Jan R Peters, and Gerhard Neumann · 2013
Earlier work this paper cites.
Probabilistic segmentation applied to an assembly task
Rudolf Lioutikov, Gerhard Neumann, Guilherme Maeda, and Jan Peters · 2015
Earlier work this paper cites.
Actor-mimic: Deep multitask and transfer reinforcement learning
Emilio Parisotto, Jimmy Lei Ba, and Ruslan Salakhutdinov · 2015
Earlier work this paper cites.
Nonparametric bayesian reward segmentation for skill discovery using inverse reinforcement learning
Pravesh Ranchod, Benjamin Rosman, and George Konidaris · 2015
Earlier work this paper cites.
Andrei A Rusu, Sergio Gomez Colmenarejo, Caglar Gulcehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, and Raia Hadsell · 2015
Earlier work this paper cites.
Learning and transfer of modulated locomotor controllers
Nicolas Heess, Greg Wayne, Yuval Tassa, Timothy Lillicrap, Martin Riedmiller, and David Silver · 2016
Earlier work this paper cites.
Efficient unsupervised temporal segmentation of motion data
Björn Krüger, Anna Vögele, Tobias Willig, Angela Yao, Reinhard Klein, and Andreas Weber · 2016
Earlier work this paper cites.
The option-critic architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup · 2017
Earlier work this paper cites.
Stochastic neural networks for hierarchical reinforcement learning
Carlos Florensa, Yan Duan, and Pieter Abbeel · 2017
Earlier work this paper cites.
Learning movement primitive libraries through probabilistic segmentation
Rudolf Lioutikov, Gerhard Neumann, Guilherme Maeda, and Jan Peters · 2017
Cited alongside, same era.
Maximum a posteriori policy optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, and Martin Riedmiller · 2018
Cited alongside, same era.
Latent space policies for hierarchical reinforcement learning
Tuomas Haarnoja, Kristian Hartikainen, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Learning an embedding space for transferable robot skills
Karol Hausman, Jost Tobias Springenberg, Ziyu Wang, Nicolas Heess, and Martin Riedmiller · 2018
Cited alongside, same era.
Using probabilistic movement primitives in robotics
Alexandros Paraschos, Christian Daniel, Jan Peters, and Gerhard Neumann · 2018
Cited alongside, same era.
Learning by playing solving sparse reward tasks from scratch
Martin Riedmiller, Roland Hafner, Thomas Lampe, Michael Neunert, Jonas Degrave, Tom van de Wiele, Vlad Mnih, Nicolas Heess, and Jost Tobias Springenberg · 2018
Advantage weighted regression: Simple and scalable off-policy reinforcement learning, 2020
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2020
Later among the works it cites.
Accelerating reinforcement learning with learned skill priors
Karl Pertsch, Youngwoon Lee, and Joseph J Lim · 2020
Later among the works it cites.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
Noah Y Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki, Michael Neunert, Thomas Lampe, Roland Hafner, Nicolas Heess, and Martin Riedmiller · 2020
Later among the works it cites.
Behavior priors for efficient reinforcement learning
D Tirumala, A Galashov, H Noh, L Hasenclever, and others · 2020
Later among the works it cites.
Critic regularized regression
Ziyu Wang, Alexander Novikov, Konrad Zolna, Josh S Merel, Jost Tobias Springenberg, Scott E Reed, Bobak Shahriari, Noah Siegel, Caglar Gulcehre, Nicolas Heess, et al · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Kickstarting deep reinforcement learning
Simon Schmitt, Jonathan J. Hudson, Augustin Zídek, Simon Osindero, Carl Doersch, Wojciech M. Czarnecki, Joel Z. Leibo, Heinrich Küttler, Andrew Zisserman, Karen Simonyan, and S. M. Ali Eslami · 2018
Cited alongside, same era.
Taco: Learning task decomposition via temporal alignment for control
Kyriacos Shiarlis, Markus Wulfmeier, Sasha Salter, Shimon Whiteson, and Ingmar Posner · 2018
Cited alongside, same era.
Information asymmetry in KL-regularized RL
Alexandre Galashov, Siddhant Jayakumar, Leonard Hasenclever, Dhruva Tirumala, Jonathan Schwarz, Guillaume Desjardins, Wojtek M Czarnecki, Yee Whye Teh, Razvan Pascanu, and Nicolas Heess · 2019
Cited alongside, same era.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind Picard · 2019
Cited alongside, same era.
Compile: Compositional imitation learning and execution
Thomas Kipf, Yujia Li, Hanjun Dai, Vinicius Zambaldi, Alvaro Sanchez-Gonzalez, Edward Grefenstette, Pushmeet Kohli, and Peter Battaglia · 2019
Cited alongside, same era.
Sub-policy adaptation for hierarchical reinforcement learning
Alexander C Li, Carlos Florensa, Ignasi Clavera, and Pieter Abbeel · 2019
Cited alongside, same era.
Later among the works it cites.
Compositional transfer in hierarchical reinforcement learning
Markus Wulfmeier, Abbas Abdolmaleki, Roland Hafner, Jost Tobias Springenberg, Michael Neunert, Tim Hertweck, Thomas Lampe, Noah Siegel, Nicolas Heess, and Martin A Riedmiller · 2020
Later among the works it cites.
Hierarchical reinforcement learning by discovering intrinsic options
Jesse Zhang, Haonan Yu, and Wei Xu · 2020
Later among the works it cites.
On multi-objective policy optimization as a tool for reinforcement learning
Abbas Abdolmaleki, Sandy H Huang, Giulia Vezzani, Bobak Shahriari, Jost Tobias Springenberg, Shruti Mishra, Dhruva TB, Arunkumar Byravan, Konstantinos Bousmalis, Andras Gyorgy, et al · 2021
Later among the works it cites.
OPAL: Offline primitive discovery for accelerating offline reinforcement learning
Anurag Ajay, Aviral Kumar, Pulkit Agrawal, Sergey Levine, and Ofir Nachum · 2021
Later among the works it cites.
Evaluating model-based planning and planner amortization for continuous control, 2021
Arunkumar Byravan, Leonard Hasenclever, Piotr Trochim, Mehdi Mirza, Alessandro Davide Ialongo, Yuval Tassa, Jost Tobias Springenberg, Abbas Abdolmaleki, Nicolas Heess, Josh Merel, and Martin Riedmiller · 2021
Later among the works it cites.
Beyond fine-tuning: Transferring behavior in reinforcement learning
Víctor Campos, Pablo Sprechmann, Steven Stenberg Hansen, André Barreto, Steven Kapturowski, Alex Vitvitskyi, Adrià Puigdomènech Badia, and Charles Blundell · 2021
Later among the works it cites.
From motor control to team play in simulated humanoid football, 2021
Siqi Liu, Guy Lever, Zhe Wang, Josh Merel, S. M. Ali Eslami, Daniel Hennes, Wojciech M. Czarnecki, Yuval Tassa, Shayegan Omidshafiei, Abbas Abdolmaleki, Noah Y. Siegel, Leonard Hasenclever, Luke Marris, Saran Tunyasuvunakool, H. Francis Song, Markus Wulfmeier, Paul Muller, Tuomas Haarnoja, Brendan D. Tracey, Karl Tuyls, Thore Graepel, and Nicolas Heess · 2021
Later among the works it cites.
Bayesian controller fusion: Leveraging control priors in deep reinforcement learning for robotics
Krishan Rana, Vibhavari Dasagi, Jesse Haviland, Ben Talbot, Michael Milford, and Niko Sünderhauf · 2021
Later among the works it cites.
Parrot: Data-Driven behavioral priors for reinforcement learning
Avi Singh, Huihan Liu, Gaoyue Zhou, Albert Yu, Nicholas Rhinehart, and Sergey Levine · 2021
Later among the works it cites.
Skid raw: Skill discovery from raw trajectories
Daniel Tanneberg, Kai Ploeger, Elmar Rueckert, and Jan Peters · 2021
Later among the works it cites.
Data-efficient hindsight off-policy option learning
Markus Wulfmeier, Dushyant Rao, Roland Hafner, Thomas Lampe, Abbas Abdolmaleki, Tim Hertweck, Michael Neunert, Dhruva Tirumala, Noah Siegel, Nicolas Heess, et al · 2021
Later among the works it cites.
Imitate and repurpose: Learning reusable robot movement skills from human and animal behaviors
Steven Bohez, Saran Tunyasuvunakool, Philemon Brakel, Fereshteh Sadeghi, Leonard Hasenclever, Yuval Tassa, Emilio Parisotto, Jan Humplik, Tuomas Haarnoja, Roland Hafner, et al · 2022
Closest in time.
The primacy bias in deep reinforcement learning
Evgenii Nikishin, Max Schwarzer, Pierluca D’Oro, Pierre-Luc Bacon, and Aaron Courville · 2022
Closest in time.
Learning transferable motor skills with hierarchical latent mixture policies
Dushyant Rao, Fereshteh Sadeghi, Leonard Hasenclever, Markus Wulfmeier, Martina Zambelli, Giulia Vezzani, Dhruva Tirumala, Yusuf Aytar, Josh Merel, Nicolas Heess, and Raia Hadsell · 2022
Closest in time.
Collect & infer-a fresh look at data-efficient reinforcement learning
Martin Riedmiller, Jost Tobias Springenberg, Roland Hafner, and Nicolas Heess · 2022
Closest in time.
Mo2: Model-based offline options
Sasha Salter, Markus Wulfmeier, Dhruva Tirumala, Nicolas Manfred Otto Heess, Martin A. Riedmiller, Raia Hadsell, and Dushyant Rao · 2022
Closest in time.
Strength through diversity: Robust behavior learning via mixture policies
Tim Seyde, Wilko Schwarting, Igor Gilitschenski, Markus Wulfmeier, and Daniela Rus · 2022
Closest in time.
Bobak Shahriari, Abbas Abdolmaleki, Arunkumar Byravan, Abe Friesen, Siqi Liu, Jost Tobias Springenberg, Nicolas Heess, Matt Hoffman, and Martin Riedmiller · 2022
Closest in time.
NeRFSim: Sim2real transfer of vision-guided bipedal motion skills using neural radiance fields
Arunkumar Byravan, Jan Humplik, Leonard Hasenclever, Arthur Brussee, Francesco Nori, Tuomas Haarnoja, Ben Moran, Fereshteh Sadeghi, Bojan Vujatovic, and Nicolas Heess · 2023
Closest in time.