D4rl: Datasets for deep data-driven reinforcement learning
Original
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Reward-rational (implicit) choice: A unifying formalism for reward learning
Hong Jun Jeon, Smitha Milli, and Anca D Dragan · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
Original
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Original
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Later among the works it cites.
Learning human objectives by evaluating hypothetical behavior
Siddharth Reddy, Anca Dragan, Sergey Levine, Shane Legg, and Jan Leike · 2020
Later among the works it cites.
The magical benchmark for robust imitation
Original
Sam Toyer, Rohin Shah, Andrew Critch, and Stuart Russell · 2020
Later among the works it cites.
Offline learning from demonstrations and unlabeled experience
Original
Konrad Zolna, Alexander Novikov, Ksenia Konyushkova, Caglar Gulcehre, Ziyu Wang, Yusuf Aytar, Misha Denil, Nando de Freitas, and Scott Reed · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron Courville, and Marc G. Bellemare · 2021
Later among the works it cites.
A survey of inverse reinforcement learning: Challenges, methods and progress
Saurabh Arora and Prashant Doshi · 2021
Later among the works it cites.
Learning" what-if" explanations for sequential decision-making
Ioana Bica, Daniel Jarrett, Alihan Hüyük, and Mihaela van der Schaar · 2021
Later among the works it cites.
Replacing rewards with examples: Example-based policy search via recursive classification
Original
Benjamin Eysenbach, Sergey Levine, and Ruslan Salakhutdinov · 2021
Later among the works it cites.
Iq-learn: Inverse soft-q learning for imitation
Divyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song, and Stefano Ermon · 2021
Later among the works it cites.
Thriftydagger: Budget-aware novelty and risk gating for interactive imitation learning
Ryan Hoque, Ashwin Balakrishna, Ellen Novoseller, Albert Wilcox, Daniel S Brown, and Ken Goldberg · 2021
Later among the works it cites.
Policy gradient bayesian robust optimization for imitation learning
Zaynah Javed, Daniel S Brown, Satvik Sharma, Jerry Zhu, Ashwin Balakrishna, Marek Petrik, Anca Dragan, and Ken Goldberg · 2021
Later among the works it cites.
Demodice: Offline imitation learning with supplementary imperfect demonstrations
Geon-Hyeong Kim, Seokin Seo, Jongmin Lee, Wonseok Jeon, HyeongJoo Hwang, Hongseok Yang, and Kee-Eung Kim · 2021
Later among the works it cites.
Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsupervised pre-training
Original
Kimin Lee, Laura Smith, and Pieter Abbeel · 2021
Later among the works it cites.
Self-supervised online reward shaping in sparse-reward environments
Farzan Memarian, Wonjoon Goo, Rudolf Lioutikov, Ufuk Topcu, and Scott Niekum · 2021
Later among the works it cites.