Fetching the paper…
Reading the bibliography…
Alvinn: An autonomous land vehicle in a neural network
D. A. Pomerleau · 1988
Earlier work this paper cites.
Batch Reinforcement Learning
S. Lange, T. Gabel, and M. A. Riedmiller · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
Trust Region Policy Optimization
J. Schulman, S. Levine, P. Moritz, M. I. Jordan, and P. Abbeel · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Earlier work this paper cites.
Proximal Policy Optimization Algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
S. Gu, E. Holly, T. Lillicrap, and S. Levine · 2017
Earlier work this paper cites.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, T. P. Lillicrap, K. Simonyan, and D. Hassabis · 2018
Earlier work this paper cites.
Off-Policy Deep Reinforcement Learning without Exploration
S. Fujimoto, D. Meger, and D. Precup · 2018
Earlier work this paper cites.
An algorithmic perspective on imitation learning
T. Osa, J. Pajarinen, G. Neumann, J. A. Bagnell, P. Abbeel, J. Peters, et al · 2018
Earlier work this paper cites.
Deep imitation learning for complex manipulation tasks from virtual reality teleoperation
T. Zhang, Z. McCarthy, O. Jow, D. Lee, X. Chen, K. Goldberg, and P. Abbeel · 2018
Earlier work this paper cites.
Truncated horizon policy search: Combining reinforcement learning & imitation learning
W. Sun, J. A. Bagnell, and B. Boots · 2018
Earlier work this paper cites.
When to trust your model: Model-based policy optimization
M. Janner, J. Fu, M. Zhang, and S. Levine · 2019
Earlier work this paper cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. J. Dudzik, J. Chung, D. H. Choi, R. W. Powell, T. Ewalds, P. Georgiev, J. Oh, D. Horgan, M. Kroiss, I. Danihelka, A. Huang, L. Sifre, T. Cai, J. P. Agapiou, M. Jaderberg, A. S. Vezhnevets, R. Leblond, T. Pohlen, V. Dalibard, D. Budden, Y. Sulsky, J. Molloy, T. L. Paine, C. Gulcehre, Z. Wang, T. Pfaff, Y. Wu, R. Ring, D. Yogatama, D. Wünsch, K. McKinney, O. Smith, T. Schaul, T. P. Lillicrap, K. Kavukcuoglu, D. Hassabis, C. Apps, and D. Silver · 2019
Earlier work this paper cites.
Behavior Regularized Offline Reinforcement Learning
Y. Wu, G. Tucker, and O. Nachum · 2019
Earlier work this paper cites.
Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
N. Jaques, A. Ghandeharioun, J. H. Shen, C. Ferguson, À. Lapedriza, N. Jones, S. Gu, and R. W. Picard · 2019
Earlier work this paper cites.
Sqil: Imitation learning via reinforcement learning with sparse rewards
S. Reddy, A. D. Dragan, and S. Levine · 2019
Earlier work this paper cites.
A Game Theoretic Framework for Model-Based Reinforcement Learning
A. Rajeswaran, I. Mordatch, and V. Kumar · 2020
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Earlier work this paper cites.
MOReL : Model-Based Offline Reinforcement Learning
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims · 2020
Cited alongside, same era.
Conservative Q-Learning for Offline Reinforcement Learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Cited alongside, same era.
Critic regularized regression
Z. Wang, A. Novikov, K. Zolna, J. S. Merel, J. T. Springenberg, S. E. Reed, B. Shahriari, N. Siegel, C. Gulcehre, N. Heess, et al · 2020
Cited alongside, same era.
Plas: Latent action space for offline reinforcement learning
W. Zhou, S. Bajracharya, and D. Held · 2020
Cited alongside, same era.
Cog: Connecting new skills to past experience with offline reinforcement learning
A. Singh, A. Yu, J. Yang, J. Zhang, A. Kumar, and S. Levine · 2020
Cited alongside, same era.
Mt-opt: Continuous multi-task robotic reinforcement learning at scale
D. Kalashnikov, J. Varley, Y. Chebotar, B. Swanson, R. Jonschkowski, C. Finn, S. Levine, and K. Hausman · 2021
Later among the works it cites.
Offline reinforcement learning from images with latent space models
R. Rafailov, T. Yu, A. Rajeswaran, and C. Finn · 2021
Later among the works it cites.
A workflow for offline model-free robotic reinforcement learning
A. Kumar, A. Singh, S. Tian, C. Finn, and S. Levine · 2021
Later among the works it cites.
What matters in learning from offline human demonstrations for robot manipulation
A. Mandlekar, D. Xu, J. Wong, S. Nasiriany, C. Wang, R. Kulkarni, L. Fei-Fei, S. Savarese, Y. Zhu, and R. Martín-Martín · 2021
Later among the works it cites.
Offline inverse reinforcement learning
F. Jarboui and V. Perchet · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Accelerating online reinforcement learning with offline datasets
A. Nair, M. Dalal, A. Gupta, and S. Levine · 2020
Cited alongside, same era.
D4RL: Datasets for Deep Data-Driven Reinforcement Learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Cited alongside, same era.
Rl unplugged: Benchmarks for offline reinforcement learning
C. Gulcehre, Z. Wang, A. Novikov, T. L. Paine, S. G. Colmenarejo, K. Zolna, R. Agarwal, J. Merel, D. Mankowitz, C. Paduraru, et al · 2020
Cited alongside, same era.
Mopo: Model-based offline policy optimization
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Y. Zou, S. Levine, C. Finn, and T. Ma · 2020
Cited alongside, same era.
Offline learning from demonstrations and unlabeled experience
K. Zolna, A. Novikov, K. Konyushkova, C. Gulcehre, Z. Wang, Y. Aytar, M. Denil, N. de Freitas, and S. Reed · 2020
Cited alongside, same era.
Semi-supervised reward learning for offline reinforcement learning
K. Konyushkova, K. Zolna, Y. Aytar, A. Novikov, S. Reed, S. Cabi, and N. de Freitas · 2020
Cited alongside, same era.
Decision Transformer: Reinforcement Learning via Sequence Modeling
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch · 2021
Cited alongside, same era.
Later among the works it cites.
Grasping with chopsticks: Combating covariate shift in model-free imitation learning for fine manipulation
L. Ke, J. Wang, T. Bhattacharjee, B. Boots, and S. Srinivasa · 2021
Later among the works it cites.
Mitigating covariate shift in imitation learning via offline data without great coverage
J. Chang, M. Uehara, D. Sreenivas, R. Kidambi, and W. Sun · 2021
Later among the works it cites.
Visual adversarial imitation learning using variational models
R. Rafailov, T. Yu, A. Rajeswaran, and C. Finn · 2021
Later among the works it cites.
A review of robot learning for manipulation: Challenges, representations, and algorithms
O. Kroemer, S. Niekum, and G. D. Konidaris · 2021
Later among the works it cites.
Cliport: What and where pathways for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2021
Later among the works it cites.
STORM: An integrated framework for fast joint-space model-predictive control for reactive manipulation
M. Bhardwaj, B. Sundaralingam, A. Mousavian, N. D. Ratliff, D. Fox, F. Ramos, and B. Boots · 2021
Later among the works it cites.
d3rlpy: An offline deep reinforcement library
M. I. Takuma Seno · 2021
Later among the works it cites.
Offline meta-reinforcement learning for industrial insertion
T. Z. Zhao, J. Luo, O. Sushkov, R. Pevceviciute, N. Heess, J. Scholz, S. Schaal, and S. Levine · 2022
Closest in time.
Don’t change the algorithm, change the data: Exploratory data for offline reinforcement learning
D. Yarats, D. Brandfonbrener, H. Liu, M. Laskin, P. Abbeel, A. Lazaric, and L. Pinto · 2022
Closest in time.
When should we prefer offline reinforcement learning over behavioral cloning?
A. Kumar, J. Hong, A. Singh, and S. Levine · 2022
Closest in time.
When should we prefer offline reinforcement learning over behavioral cloning?
A. Kumar, J. Hong, A. Singh, and S. Levine · 2022
Closest in time.
Implicit behavioral cloning
P. Florence, C. Lynch, A. Zeng, O. A. Ramirez, A. Wahid, L. Downs, A. Wong, J. Lee, I. Mordatch, and J. Tompson · 2022
Closest in time.
R3m: A universal visual representation for robot manipulation
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Closest in time.
A survey on offline reinforcement learning: Taxonomy, review, and open problems
R. F. Prudencio, M. R. Maximo, and E. L. Colombini · 2022
Closest in time.