Fetching the paper…
Reading the bibliography…
This work introduces Transformer-based Off-Policy Episodic Reinforcement Learning (TOP-ERL), a novel algorithm that enables off-policy updates in the ERL framework.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Genetic reinforcement learning for neurocontrol problems
Darrell Whitley, Stephen Dominic, Rajarshi Das, and Charles W Anderson · 1993
Earlier work this paper cites.
On generating power law noise
Jens Timmer and Michel Koenig · 1995
Earlier work this paper cites.
Neuroevolution for reinforcement learning using evolution strategies
Christian Igel · 2003
Earlier work this paper cites.
Dynamic movement primitives-a framework for motor control in humans and humanoid robotics
Stefan Schaal · 2006
Earlier work this paper cites.
Accelerated neural evolution through cooperatively coevolved synapses
Faustino Gomez, Jürgen Schmidhuber, Risto Miikkulainen, and Melanie Mitchell · 2008
Earlier work this paper cites.
Policy search for motor primitives in robotics
Jens Kober and Jan Peters · 2008
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
Jan Peters and Stefan Schaal · 2008
Earlier work this paper cites.
State-dependent exploration for policy gradient methods
Thomas Rückstieß, Martin Felder, and Jürgen Schmidhuber · 2008
Earlier work this paper cites.
Exploring parameter space in reinforcement learning
Thomas Rückstiess, Frank Sehnke, Tom Schaul, Daan Wierstra, Yi Sun, and Jürgen Schmidhuber · 2010
Earlier work this paper cites.
Dynamical movement primitives: learning attractor models for motor behaviors
Auke Jan Ijspeert, Jun Nakanishi, Heiko Hoffmann, Peter Pastor, and Stefan Schaal · 2013
Earlier work this paper cites.
Probabilistic movement primitives
Alexandros Paraschos, Christian Daniel, Jan Peters, and Gerhard Neumann · 2013
Earlier work this paper cites.
Learning interaction for collaborative tasks with probabilistic movement primitives
Guilherme Maeda, Marco Ewerton, Rudolf Lioutikov, Heni Ben Amor, Jan Peters, and Gerhard Neumann · 2014
Earlier work this paper cites.
Model-based relative entropy stochastic search
Abbas Abdolmaleki, Rudolf Lioutikov, Jan R Peters, Nuno Lau, Luis Pualo Reis, and Gerhard Neumann · 2015
Earlier work this paper cites.
Frame skip is a powerful parameter for learning to play atari
Alex Braylan, Mark Hollenbeck, Elliot Meyerson, and Risto Miikkulainen · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Jimmy Lei Ba · 2016
Earlier work this paper cites.
Using probabilistic movement primitives for striking movements
Sebastian Gomez-Gonzalez, Gerhard Neumann, Bernhard Schölkopf, and Jan Peters · 2016
Earlier work this paper cites.
Contextual covariance matrix adaptation evolutionary strategies
Abbas Abdolmaleki, Bob Price, Nuno Lau, Luis Paulo Reis, and Gerhard Neumann · 2017
Earlier work this paper cites.
Evolution strategies as a scalable alternative to reinforcement learning
Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Attention is all you need
A Vaswani · 2017
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Cited alongside, same era.
Rllib: Abstractions for distributed reinforcement learning
Eric Liang, Richard Liaw, Robert Nishihara, Philipp Moritz, Roy Fox, Ken Goldberg, Joseph Gonzalez, Michael Jordan, and Ion Stoica · 2018
Cited alongside, same era.
Simple random search of static linear policies is competitive for reinforcement learning
Horia Mania, Aurelia Guy, and Benjamin Recht · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Differentiable convex optimization layers
Akshay Agrawal, Brandon Amos, Shane Barratt, Stephen Boyd, Steven Diamond, and J Zico Kolter · 2019
Cited alongside, same era.
Smooth exploration for robotic reinforcement learning
Antonin Raffin, Jens Kober, and Freek Stulp · 2022
Later among the works it cites.
Orientation probabilistic movement primitives on riemannian manifolds
Leonel Rozo and Vedant Dave · 2022
Later among the works it cites.
Online decision transformer
Qinqing Zheng, Amy Zhang, and Aditya Grover · 2022
Later among the works it cites.
Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions
Yevgen Chebotar, Quan Vuong, Karol Hausman, Fei Xia, Yao Lu, Alex Irpan, Aviral Kumar, Tianhe Yu, Alexander Herzog, Karl Pertsch, et al · 2023
Later among the works it cites.
Prodmp: A unified perspective on dynamic and probabilistic movement primitives
Ge Li, Zeqi Jin, Michael Volpp, Fabian Otto, Rudolf Lioutikov, and Gerhard Neumann · 2023
Later among the works it cites.
Model-based reinforcement learning with multi-step plan value estimation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Improving local trajectory optimisation using probabilistic movement primitives
RB Ashith Shyam, Peter Lightbody, Gautham Das, Pengcheng Liu, Sebastian Gomez-Gonzalez, and Gerhard Neumann · 2019
Cited alongside, same era.
Learning via-point movement primitives with inter-and extrapolation capabilities
You Zhou, Jianfeng Gao, and Tamim Asfour · 2019
Cited alongside, same era.
Neural dynamic policies for end-to-end sensorimotor learning
Shikhar Bahl, Mustafa Mukadam, Abhinav Gupta, and Deepak Pathak · 2020
Cited alongside, same era.
Training of deep neural networks for the generation of dynamic movement primitives
Rok Pahič, Barry Ridge, Andrej Gams, Jun Morimoto, and Aleš Ude · 2020
Cited alongside, same era.
Stabilizing transformers for reinforcement learning
Emilio Parisotto, Francis Song, Jack Rae, Razvan Pascanu, Caglar Gulcehre, Siddhant Jayakumar, Max Jaderberg, Raphael Lopez Kaufman, Aidan Clark, Seb Noury, et al · 2020
Cited alongside, same era.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Karol Hausman, Chelsea Finn, and Sergey Levine · 2020
Cited alongside, same era.
Haoxin Lin, Yihao Sun, Jiaji Zhang, and Yang Yu · 2023
Later among the works it cites.
Mp3: Movement primitive-based (re-) planning policy
Fabian Otto, Hongyi Zhou, Onur Celik, Ge Li, Rudolf Lioutikov, and Gerhard Neumann · 2023
Later among the works it cites.
Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl
Taku Yamagata, Ahmed Khalil, and Raul Santos-Rodriguez · 2023
Later among the works it cites.
Policy expansion for bridging offline-to-online reinforcement learning
Haichao Zhang, We Xu, and Haonan Yu · 2023
Later among the works it cites.
Learning fine-grained bimanual manipulation with low-cost hardware
Tony Z Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn · 2023
Later among the works it cites.
Roboagent: Generalization and efficiency in robot manipulation via semantic augmentations and action chunking
Homanga Bharadhwaj, Jay Vakil, Mohit Sharma, Abhinav Gupta, Shubham Tulsiani, and Vikash Kumar · 2024
Closest in time.
Acquiring diverse skills using curriculum reinforcement learning with mixture of experts
Onur Celik, Aleksandar Taranovic, and Gerhard Neumann · 2024
Closest in time.
Bridging the gap between learning-to-plan, motion primitives and safe reinforcement learning
Piotr Kicki, Davide Tateo, Puze Liu, Jonas Günster, Jan Peters, and Krzysztof Walas · 2024
Closest in time.
Open the black box: Step-based policy updates for temporally-correlated episodic reinforcement learning
Ge Li, Hongyi Zhou, Dominik Roth, Serge Thilges, Fabian Otto, Rudolf Lioutikov, and Gerhard Neumann · 2024
Closest in time.
Rethinking transformers in solving POMDPs
Chenhao Lu, Ruizhe Shi, Yuyao Liu, Kaizhe Hu, Simon Shaolei Du, and Huazhe Xu · 2024
Closest in time.
Weighting online decision transformer with episodic memory for offline-to-online reinforcement learning
Xiao Ma and Wu-Jun Li · 2024
Closest in time.
When do transformers shine in rl? decoupling memory from credit assignment
Tianwei Ni, Michel Ma, Benjamin Eysenbach, and Pierre-Luc Bacon · 2024
Closest in time.
Decision mamba: Reinforcement learning via sequence modeling with selective state spaces
Toshihiro Ota · 2024
Closest in time.
Multimodal diffusion transformer: Learning versatile behavior from multimodal goals
Moritz Reuss, Ömer Erdinç Yağmurlu, Fabian Wenzel, and Rudolf Lioutikov · 2024
Closest in time.
Elastic decision transformer
Yueh-Hua Wu, Xiaolong Wang, and Masashi Hamaya · 2024
Closest in time.
Transformer in reinforcement learning for decision-making: a survey
Weilin Yuan, Jiaxing Chen, Shaofei Chen, Dawei Feng, Zhenzhen Hu, Peng Li, and Weiwei Zhao · 2024
Closest in time.