Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (RL) algorithms can learn better decision-making compared to behavior policies by stitching the suboptimal trajectories to derive more optimal ones.
A survey of robot learning from demonstration
Brenna D. Argall, Sonia Chernova, Manuela Veloso, and Brett Browning · 2008
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stephane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Imitation learning with demonstrations and shaping rewards
Kshitij Judah, Alan Fern, Prasad Tadepalli, and Robby Goetschalckx · 2014
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Imitation from observation: Learning to imitate behaviors from raw video via context translation
YuXuan Liu, Abhishek Gupta, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Diagnosing bottlenecks in deep q-learning algorithms
Justin Fu, Aviral Kumar, Matthew Soh, and Sergey Levine · 2019
Earlier work this paper cites.
Imitation learning via off-policy distribution matching
Ilya Kostrikov, Ofir Nachum, and Jonathan Tompson · 2019
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2019
Earlier work this paper cites.
Sqil: Imitation learning via reinforcement learning with sparse rewards
Siddharth Reddy, Anca D. Dragan, and Sergey Levine · 2019
Earlier work this paper cites.
Recent advances in imitation learning from observation
Faraz Torabi, Garrett Warnell, and Peter Stone · 2019
Earlier work this paper cites.
Disagreement-regularized imitation learning
Kiante Brantley, Wen Sun, and Mikael Henaff · 2020
Earlier work this paper cites.
Better-than-demonstrator imitation learning via automatically-ranked demonstrations
Daniel S. Brown, Wonjoon Goo, and Scott Niekum · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, G. Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Cited alongside, same era.
Recent advances in robot learning from demonstration
Harish Ravichandar, Athanasios S. Polydoros, Sonia Chernova, and Aude Billard · 2020
Iq-learn: Inverse soft-q learning for imitation
Divyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song, Matthieu Geist, and Stefano Ermon · 2022
Later among the works it cites.
DemoDICE: Offline imitation learning with supplementary imperfect demonstrations
Geon-Hyeong Kim, Seokin Seo, Jongmin Lee, Wonseok Jeon, HyeongJoo Hwang, Hongseok Yang, and Kee-Eung Kim · 2022
Later among the works it cites.
Versatile offline imitation from observations and examples via regularized state-occupancy matching
Yecheng Jason Ma, Andrew Shen, Dinesh Jayaraman, and Osbert Bastani · 2022
Later among the works it cites.
You can’t count on luck: Why decision transformers and rvs fail in stochastic environments
Keiran Paster, Sheila McIlraith, and Jimmy Ba · 2022
Later among the works it cites.
Supported policy optimization for offline reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Offline learning from demonstrations and unlabeled experience
Konrad Zolna, Alexander Novikov, Ksenia Konyushkova, Caglar Gulcehre, Ziyu Wang, Yusuf Aytar, Misha Denil, Nando de Freitas, and Scott Reed · 2020
Cited alongside, same era.
Uncertainty-based offline reinforcement learning with diversified q-ensemble
Gaon An, Seungyong Moon, Jang-Hyun Kim, and Hyun Oh Song · 2021
Cited alongside, same era.
Mitigating covariate shift in imitation learning via offline data with partial coverage
Jonathan Chang, Masatoshi Uehara, Dhruv Sreenivas, Rahul Kidambi, and Wen Sun · 2021
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch · 2021
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2021
Cited alongside, same era.
Offline reinforcement learning with implicit q-learning
Ilya Kostrikov, Ashvin Nair, and Sergey Levine · 2021
Cited alongside, same era.
Behavioral cloning from noisy demonstrations
Fumihiro Sasaki and Ryota Yamashina · 2021
Cited alongside, same era.
Jialong Wu, Haixu Wu, Zihan Qiu, Jianmin Wang, and Mingsheng Long · 2022
Later among the works it cites.
Prompting decision transformer for few-shot policy generalization
Mengdi Xu, Yikang Shen, Shun Zhang, Yuchen Lu, Ding Zhao, Joshua B. Tenenbaum, and Chuang Gan · 2022
Later among the works it cites.
Ditto: Offline imitation learning with world models
Branton DeMoss, Paul Duckworth, Nick Hawes, and Ingmar Posner · 2023
Later among the works it cites.
Pre-training to learn in context
Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang · 2023
Later among the works it cites.
Offline rl with observation histories: Analyzing and improving sample complexity
Joey Hong, Anca Dragan, and Sergey Levine · 2023
Later among the works it cites.
Prompt-tuning decision transformer with preference ranking
Shengchao Hu, Li Shen, Ya Zhang, and Dacheng Tao · 2023
Later among the works it cites.
Beyond reward: Offline preference-guided policy optimization
Yachen Kang, Diyuan Shi, Jinxin Liu, Li He, and Donglin Wang · 2023
Later among the works it cites.
Ceil: Generalized contextual imitation learning
Jinxin Liu, Li He, Yachen Kang, Zifeng Zhuang, Donglin Wang, and Huazhe Xu · 2023
Later among the works it cites.
Taku Yamagata, Ahmed Khalil, and Raul Santos-Rodriguez · 2023
Later among the works it cites.
Discriminator-guided model-based offline imitation learning
Wenjia Zhang, Haoran Xu, Haoyi Niu, Peng Cheng, Ming Li, Heming Zhang, Guyue Zhou, and Xianyuan Zhan · 2023
Later among the works it cites.