Fetching the paper…
Reading the bibliography…
Pre-training with offline data and online fine-tuning using reinforcement learning is a promising strategy for learning control policies by leveraging the best of both worlds in terms of sample efficiency and performance.
ALVINN: An autonomous land vehicle in a neural network
Dean A. Pomerleau · 1988
Earlier work this paper cites.
A framework for behavioural cloning
Michael Bain and Claude Sammut · 1996
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. Mcallester, S. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
Off-road obstacle avoidance through end-to-end learning
Urs Muller, Jan Ben, Eric Cosatto, Beat Flepp, and Yann LeCun · 2005
Earlier work this paper cites.
Visualizing high-dimensional data using t-SNE
Laurens van der Maaten and Geoffrey E. Hinton · 2008
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stephane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Agnostic system identification for model-based reinforcement learning
Stephane Ross and Drew Bagnell · 2012
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Jens Kober, J. Andrew Bagnell, and Jan Peters · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Earlier work this paper cites.
Policy distillation
Andrei A. Rusu, Sergio Gomez Colmenarejo, Çaglar Gülçehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, and Raia Hadsell · 2016
Earlier work this paper cites.
Borrowing treasures from the wealthy: Deep transfer learning through selective joint fine-tuning
Weifeng Ge and Yizhou Yu · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Mastering the game of Go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy P. Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis · 2017
Earlier work this paper cites.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
Matej Vecerík, Todd Hester, Jonathan Scholz, Fumin Wang, Olivier Pietquin, Bilal Piot, Nicolas Heess, Thomas Rothörl, Thomas Lampe, and Martin A. Riedmiller · 2017
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Deep Q-learning from demonstrations
Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Dan Horgan, John Quan, Andrew Sendonaris, Ian Osband, Gabriel Dulac-Arnold, John Agapiou, Joel Z. Leibo, and Audrunas Gruslys · 2018
Earlier work this paper cites.
Overcoming exploration in reinforcement learning with demonstrations
Ashvin Nair, Bob McGrew, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford and Karthik Narasimhan · 2018
Earlier work this paper cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
Aravind Rajeswaran, Vikash Kumar, Abhishek Gupta, Giulia Vezzani, John Schulman, Emanuel Todorov, and Sergey Levine · 2018
Earlier work this paper cites.
A data-driven approach for autonomous motion planning and control in off-road driving scenarios
Hossein Rastgoftar, Bingxin Zhang, and Ella M. Atkins · 2018
Cited alongside, same era.
Hierarchical and interpretable skill acquisition in multi-task reinforcement learning
Tianmin Shu, Caiming Xiong, and Richard Socher · 2018
Cited alongside, same era.
A framework for data-driven robotics
Serkan Cabi, Sergio Gómez Colmenarejo, Alexander Novikov, Ksenia Konyushkova, Scott E. Reed, Rae Jeong, Konrad Zolna, Yusuf Aytar, David Budden, Mel Vecerík, Oleg Sushkov, David Barker, Jonathan Scholz, Misha Denil, Nando de Freitas, and Ziyu Wang · 2019
Cited alongside, same era.
Exploring the limitations of behavior cloning for autonomous driving
Felipe Codevilla, Eder Santana, Antonio M. López, and Adrien Gaidon · 2019
Cited alongside, same era.
Causal confusion in imitation learning
Pim de Haan, Dinesh Jayaraman, and Sergey Levine · 2019
Cited alongside, same era.
On pathologies in KL-regularized reinforcement learning from expert demonstrations
Tim G. J. Rudner, Cong Lu, Michael Osborne, Yarin Gal, and Yee Whye Teh · 2021
Later among the works it cites.
Pretraining representations for data-efficient reinforcement learning
Max Schwarzer, Nitarshan Rajkumar, Michael Noukhovitch, Ankesh Anand, Laurent Charlin, R Devon Hjelm, Philip Bachman, and Aaron C Courville · 2021
Later among the works it cites.
RRL: Resnet as representation for reinforcement learning
Rutav Shah and Vikash Kumar · 2021
Later among the works it cites.
Parrot: Data-driven behavioral priors for reinforcement learning
Avi Singh, Huihan Liu, Gaoyue Zhou, Albert Yu, Nicholas Rhinehart, and Sergey Levine · 2021
Later among the works it cites.
Human-level reinforcement learning through theory-based modeling, exploration, and planning
Pedro A. Tsividis, João Loula, Jake Burga, Nathan Foss, Andres Campero, Thomas Pouncy, Samuel J. Gershman, and Joshua B. Tenenbaum · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Diversity is all you need: Learning skills without a reward function
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine · 2019
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Cited alongside, same era.
Do better ImageNet models transfer better?
Simon Kornblith, Jonathon Shlens, and Quoc V. Le · 2019
Cited alongside, same era.
Mastering Atari, Go, Chess and Shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, Timothy P. Lillicrap, and David Silver · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
D4RL: datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Conservative Q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Policy finetuning: Bridging sample-efficient offline and online reinforcement learning
Tengyang Xie, Nan Jiang, Huan Wang, Caiming Xiong, and Yu Bai · 2021
Later among the works it cites.
Pessimistic model selection for offline deep reinforcement learning
Chao-Han Huck Yang, Zhengling Qi, Yifan Cui, and Pin-Yu Chen · 2021
Later among the works it cites.
Representation matters: Offline pretraining for sequential decision making
Mengjiao Yang and Ofir Nachum · 2021
Later among the works it cites.
TAAC: Temporally abstract actor-critic for continuous control
Haonan Yu, Wei Xu, and Haichao Zhang · 2021
Later among the works it cites.
Offline RL policies should be trained to be adaptive
Dibya Ghosh, Anurag Ajay, Pulkit Agrawal, and Sergey Levine · 2022
Later among the works it cites.
Offline reinforcement learning with implicit Q-learning
Ilya Kostrikov, Ashvin Nair, and Sergey Levine · 2022
Later among the works it cites.
Should I run offline reinforcement learning or behavioral cloning?
Aviral Kumar, Joey Hong, Anikait Singh, and Sergey Levine · 2022
Later among the works it cites.
Challenges and opportunities in offline reinforcement learning from visual observations
Cong Lu, Philip J. Ball, Tim G. J. Rudner, Jack Parker-Holder, Michael A. Osborne, and Yee Whye Teh · 2022
Later among the works it cites.
Offline reinforcement learning as anti-exploration
Shideh Rezaeifar, Robert Dadashi, Nino Vieillard, Léonard Hussenot, Olivier Bachem, Olivier Pietquin, and Matthieu Geist · 2022
Later among the works it cites.
Reinforcement learning with action-free pre-training from videos
Younggyo Seo, Kimin Lee, Stephen L James, and Pieter Abbeel · 2022
Later among the works it cites.
Value function spaces: Skill-centric state abstractions for long-horizon reasoning
Dhruv Shah, Alexander T Toshev, Sergey Levine, and Brian Ichter · 2022
Later among the works it cites.
Jump-start reinforcement learning
Ikechukwu Uchendu, Ted Xiao, Yao Lu, Banghua Zhu, Mengyuan Yan, Joséphine Simon, Matthew Bennice, Chuyuan Fu, Cong Ma, Jiantao Jiao, Sergey Levine, and Karol Hausman · 2022
Later among the works it cites.
Discriminator-weighted offline imitation learning from suboptimal demonstrations
Haoran Xu, Xianyuan Zhan, Honglei Yin, and Huiling Qin · 2022
Later among the works it cites.
Pre-trained image encoder for generalizable visual reinforcement learning
Zhecheng Yuan, Zhengrong Xue, Bo Yuan, Xueqian Wang, Yi Wu, Yang Gao, and Huazhe Xu · 2022
Later among the works it cites.
Generative planning for temporally coordinated exploration in reinforcement learning
Haichao Zhang, Wei Xu, and Haonan Yu · 2022
Later among the works it cites.
Semi-supervised offline reinforcement learning with action-free trajectories
Qinqing Zheng, Mikael Henaff, Brandon Amos, and Aditya Grover · 2022
Later among the works it cites.