Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) provides an appealing formalism for learning control policies from experience.
Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning
Ronald J Williams · 1992
Earlier work this paper cites.
Robot Learning From Demonstration
Christopher G Atkeson and Stefan Schaal · 1997
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Actor-Critic Algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
Learning Attractor Landscapes for Learning Motor Primitives
Auke Jan Ijspeert, Jun Nakanishi, and Stefan Schaal · 2002
Earlier work this paper cites.
Learning from observation and from practice using behavioral primitives
Darrin C. Bentivegna, Gordon Cheng, and Christopher G. Atkeson · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Learning from demonstration
Stefan Schaal · 2004
Earlier work this paper cites.
Reinforcement Learning by Reward-weighted Regression for Operational Space Control
Jan Peters and Stefan Schaal · 2007
Earlier work this paper cites.
Natural actor-critic
Jan Peters and Stefan Schaal · 2007
Earlier work this paper cites.
Policy search for motor primitives in robotics
Jens Kober and J. Peter · 2008
Earlier work this paper cites.
Fitted Q-iteration by Advantage Weighted Regression
Gerhard Neumann and Jan Peters · 2008
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
Jan Peters and Stefan Schaal · 2008
Earlier work this paper cites.
Maximum Entropy Inverse Reinforcement Learning
Brian D Ziebart, Andrew Maas, J Andrew Bagnell, and Anind K Dey · 2008
Earlier work this paper cites.
Natural actor-critic algorithms
Shalabh Bhatnagar, Richard S. Sutton, Mohammad Ghavamzadeh, and Mark Lee · 2009
Earlier work this paper cites.
Real-time reinforcement learning by sequential actor-critics and experience replay
Pawel Wawrzynski · 2009
Earlier work this paper cites.
Robot motor skill coordination with em-based reinforcement learning
Petar Kormushev, Sylvain Calinon, and Darwin G. Caldwell · 2010
Earlier work this paper cites.
Relative Entropy Policy Search
Jan Peters, Katharina Mülling, and Yasemin Altün · 2010
Earlier work this paper cites.
A Generalized Path Integral Control Approach to Reinforcement Learning
Evangelos A Theodorou, Jonas Buchli, and Stefan Schaal · 2010
Earlier work this paper cites.
Off-Policy Actor-Critic
Thomas Degris, Martha White, and Richard S. Sutton · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Batch reinforcement learning
Sascha Lange, Thomas Gabel, and Martin A. Riedmiller · 2012
Cited alongside, same era.
MuJoCo: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
Learning from Limited Demonstrations
Beomjoon Kim, Amir-Massoud Farahmand, Joelle Pineau, and Doina Precup · 2013
Cited alongside, same era.
Learning to select and generalize striking movements in robot table tennis
Katharina Mülling, Jens Kober, Oliver Kroemer, and Jan Peters · 2013
Cited alongside, same era.
Compatible value gradients for reinforcement learning of continuous deep policies
David Balduzzi and Muhammad Ghifary · 2015
Cited alongside, same era.
Off-policy Model-based Learning under Unknown Factored Dynamics
Assaf Hallak, Francois Schnitzler, Timothy Mann, and Shie Mannor · 2015
Exponentially Weighted Imitation Learning for Batched Historical Data
Qing Wang, Jiechao Xiong, Lei Han, Peng Sun, Han Liu, and Tong Zhang · 2018
Later among the works it cites.
An Optimistic Perspective on Offline Reinforcement Learning
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi · 2019
Later among the works it cites.
ROBEL: Robotics Benchmarks for Learning with Low-Cost Robots
Michael Ahn, Henry Zhu, Kristian Hartikainen, Hugo Ponte, Abhishek Gupta, Sergey Levine, and Vikash Kumar · 2019
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Later among the works it cites.
P3O: Policy-on Policy-off Policy Optimization
Rasool Fakoor, Pratik Chaudhari, and Alexander J Smola · 2019
Later among the works it cites.
Off-Policy Deep Reinforcement Learning without Exploration
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Generalized Emphatic Temporal Difference Learning: Bias-Variance Analysis
Assaf Hallak, Aviv Tamar, Rémi Munos, and Shie Mannor · 2016
Cited alongside, same era.
Doubly Robust Off-policy Value Evaluation for Reinforcement Learning
Nan Jiang and Lihong Li · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Cited alongside, same era.
Asynchronous Methods for Deep Reinforcement Learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Tim Harley, Timothy P Lillicrap, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip S. Thomas and Emma Brunskill · 2016
Cited alongside, same era.
Consistent On-Line Off-Policy Evaluation
Assaf Hallak and Shie Mannor · 2017
Cited alongside, same era.
Scott Fujimoto, David Meger, and Doina Precup · 2019
Later among the works it cites.
Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning
Abhishek Gupta, Vikash Kumar, Corey Lynch, Sergey Levine, and Karol Hausman · 2019
Later among the works it cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Àgata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind W. Picard · 2019
Later among the works it cites.
Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction
Aviral Kumar, Justin Fu, George Tucker, and Sergey Levine · 2019
Later among the works it cites.
DualDICE: Behavior-Agnostic Estimation of Discounted Stationary Distribution Corrections
Ofir Nachum, Yinlam Chow, Bo Dai, and Lihong Li · 2019
Later among the works it cites.
Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2019
Later among the works it cites.
Generalized off-policy actor-critic
Shangtong Zhang, Wendelin Boehmer, and Shimon Whiteson · 2019
Later among the works it cites.
Watch, try, learn: Meta-learning from demonstrations and reward
Allan Zhou, Eric Jang, Daniel Kappler, Alexander Herzog, Mohi Khansari, Paul Wohlhart, Yunfei Bai, Mrinal Kalakrishnan, Sergey Levine, and Chelsea Finn · 2019
Later among the works it cites.
Dexterous Manipulation with Deep Reinforcement Learning: Efficient, General, and Low-Cost
Henry Zhu, Abhishek Gupta, Aravind Rajeswaran, Sergey Levine, and Vikash Kumar · 2019
Later among the works it cites.
D4RL: Datasets for Deep Data-Driven Reinforcement Learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Closest in time.
Conservative Q-Learning for Offline Reinforcement Learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Closest in time.
Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Closest in time.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning, 2020
Noah Y. Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki, Michael Neunert, Thomas Lampe, Roland Hafner, Nicolas Heess, and Martin Riedmiller · 2020
Closest in time.
Critic Regularized Regression
Ziyu Wang, Alexander Novikov, Konrad Zołna, Jost Tobias Springenberg, Scott Reed, Bobak Shahriari, Noah Siegel, Josh Merel, Caglar Gulcehre, Nicolas Heess, and Nando De Freitas · 2020
Closest in time.
GenDICE: Generalized Offline Estimation of Stationary Values
Ruiyi Zhang, Bo Dai, Lihong Li, and Dale Schuurmans · 2020
Closest in time.
Reset-Free Reinforcement Learning via Multi-Task Learning: Learning Dexterous Manipulation Behaviors without Human Intervention
Abhishek Gupta, Justin Yu, Tony Zhao, Vikash Kumar, Kelvin Xu, Thomas Devlin, Aaron Rovinsky, and Sergey Levine · 2021
Closest in time.