Fetching the paper…
Reading the bibliography…
Reinforcement Learning (RL) is notoriously data-inefficient, which makes training on a real robot difficult.
Planning and acting in partially observable stochastic domains
L. P. Kaelbling, M. L. Littman, and A. R. Cassandra · 1998
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
J. Peters and S. Schaal · 2007
Earlier work this paper cites.
Model predictive path integral control using covariance variable importance sampling
G. Williams, A. Aldrich, and E. A. Theodorou · 2015
Earlier work this paper cites.
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2015
Earlier work this paper cites.
Sim-to-real transfer of robotic control with dynamics randomization
X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel · 2017
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel · 2017
Earlier work this paper cites.
Learning dexterous in-hand manipulation
M. Andrychowicz, B. Baker, M. Chociej, R. Józefowicz, B. McGrew, J. W. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, J. Schneider, S. Sidor, J. Tobin, P. Welinder, L. Weng, and W. Zaremba · 2018
Earlier work this paper cites.
Asymmetric actor critic for image-based robot learning
L. Pinto, M. Andrychowicz, P. Welinder, W. Zaremba, and P. Abbeel · 2018
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2018
Earlier work this paper cites.
Recurrent world models facilitate policy evolution
D. Ha and J. Schmidhuber · 2018
Earlier work this paper cites.
Soft actor-critic algorithms and applications
T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V. Kumar, H. Zhu, A. Gupta, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations
A. Rajeswaran, V. Kumar, A. Gupta, G. Vezzani, J. Schulman, E. Todorov, and S. Levine · 2018
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
C. Berner, G. Brockman, B. Chan, V. Cheung, P. Debiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, R. Józefowicz, S. Gray, C. Olsson, J. W. Pachocki, M. Petrov, H. P. de Oliveira Pinto, J. Raiman, T. Salimans, J. Schlatter, J. Schneider, S. Sidor, I. Sutskever, J. Tang, F. Wolski, and S. Zhang · 2019
Earlier work this paper cites.
Mastering atari, go, chess and shogi by planning with a learned model
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, T. P. Lillicrap, and D. Silver · 2019
Earlier work this paper cites.
Behavior regularized offline reinforcement learning
Y. Wu, G. Tucker, and O. Nachum · 2019
Earlier work this paper cites.
Stabilizing off-policy Q-learning via bootstrapping error reduction
A. Kumar, J. Fu, M. Soh, G. Tucker, and S. Levine · 2019
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
X. B. Peng, A. Kumar, G. Zhang, and S. Levine · 2019
Earlier work this paper cites.
Striving for simplicity in off-policy deep reinforcement learning
R. Agarwal, D. Schuurmans, and M. Norouzi · 2019
Earlier work this paper cites.
Sim-to-real transfer in deep reinforcement learning for robotics: a survey
W. Zhao, J. P. Queralta, and T. Westerlund · 2020
Cited alongside, same era.
MOReL: Model-based offline reinforcement learning
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Cited alongside, same era.
AWAC: Accelerating online reinforcement learning with offline datasets
A. Nair, A. Gupta, M. Dalal, and S. Levine · 2020
Cited alongside, same era.
Efficient adaptation for end-to-end vision-based robotic manipulation
R. C. Julian, B. Swanson, G. S. Sukhatme, S. Levine, C. Finn, and K. Hausman · 2020
Cited alongside, same era.
Should I run offline reinforcement learning or behavioral cloning?
A. Kumar, J. Hong, A. Singh, and S. Levine · 2022
Later among the works it cites.
Video pretraining (VPT): Learning to act by watching unlabeled online videos
B. Baker, I. Akkaya, P. Zhokov, J. Huizinga, J. Tang, A. Ecoffet, B. Houghton, R. Sampedro, and J. Clune · 2022
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
I. Kostrikov, A. Nair, and S. Levine · 2022
Later among the works it cites.
Offline-to-online reinforcement learning via balanced replay and pessimistic Q-ensemble
S. Lee, Y. Seo, K. Lee, P. Abbeel, and J. Shin · 2022
Later among the works it cites.
Don’t change the algorithm, change the data: Exploratory data for offline reinforcement learning
D. Yarats, D. Brandfonbrener, H. Liu, M. Laskin, P. Abbeel, A. Lazaric, and L. Pinto · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Cited alongside, same era.
MOPO: Model-based offline policy optimization
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Y. Zou, S. Levine, C. Finn, and T. Ma · 2020
Cited alongside, same era.
A framework for efficient robotic manipulation
A. Zhan, P. Zhao, L. Pinto, P. Abbeel, and M. Laskin · 2020
Cited alongside, same era.
Self-supervised policy adaptation during deployment
N. Hansen, R. Jangir, Y. Sun, G. Alenyà, P. Abbeel, A. A. Efros, L. Pinto, and X. Wang · 2021
Cited alongside, same era.
A minimalist approach to offline reinforcement learning
S. Fujimoto and S. S. Gu · 2021
Cited alongside, same era.
Adaptive-control-oriented meta-learning for nonlinear systems
S. M. Richards, N. Azizan, J.-J. Slotine, and M. Pavone · 2021
Cited alongside, same era.
Randomized ensembled double q-learning: Learning fast without a model
X. Chen, C. Wang, Z. Zhou, and K. W. Ross · 2021
Cited alongside, same era.
Temporal difference learning for model predictive control
N. Hansen, X. Wang, and H. Su · 2022
Later among the works it cites.
Disentangling epistemic and aleatoric uncertainty in reinforcement learning
B. Charpentier, R. Senanayake, M. Kochenderfer, and S. Günnemann · 2022
Later among the works it cites.
A generalist agent
S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Gimenez, Y. Sulsky, J. Kay, J. T. Springenberg, T. Eccles, J. Bruce, A. Razavi, A. D. Edwards, N. M. O. Heess, Y. Chen, R. Hadsell, O. Vinyals, M. Bordbar, and N. de Freitas · 2022
Later among the works it cites.
Look closer: Bridging egocentric and third-person views with transformers for robotic manipulation
R. Jangir, N. Hansen, S. Ghosal, M. Jain, and X. Wang · 2022
Later among the works it cites.
Pre-training for robots: Offline RL enables learning new tasks from a handful of trials
A. Kumar, A. Singh, F. Ebert, Y. Yang, C. Finn, and S. Levine · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Gray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe · 2022
Later among the works it cites.
VRL3: A data-driven framework for visual deep reinforcement learning
C. Wang, X. Luo, K. W. Ross, and D. Li · 2022
Later among the works it cites.
Daydreamer: World models for physical robot learning
P. Wu, A. Escontrela, D. Hafner, K. Goldberg, and P. Abbeel · 2022
Later among the works it cites.
Cal-QL: Calibrated offline RL pre-training for efficient online fine-tuning
M. Nakamoto, Y. Zhai, A. Singh, M. S. Mark, Y. Ma, C. Finn, A. Kumar, and S. Levine · 2023
Closest in time.
Modem: Accelerating visual model-based reinforcement learning with demonstrations
N. Hansen, Y. Lin, H. Su, X. Wang, V. Kumar, and A. Rajeswaran · 2023
Closest in time.
Visual reinforcement learning with self-supervised 3D representations
Y. Ze, N. Hansen, Y. Chen, M. Jain, and X. Wang · 2023
Closest in time.
On the feasibility of cross-task transfer with model-based reinforcement learning
Y. Xu, N. Hansen, Z. Wang, Y.-C. Chan, H. Su, and Z. Tu · 2023
Closest in time.