Fetching the paper…
Reading the bibliography…
Future- or return-conditioned supervised learning is an emerging paradigm for offline reinforcement learning (RL), where the future outcome (i.e., return) associated with an observed action sequence is used as input to a policy trained to imitate those same actions.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Who wrote the serenity prayer?
Fred R Shapiro · 2014
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Multi-level discovery of deep options
Roy Fox, Sanjay Krishnan, Ion Stoica, and Ken Goldberg · 2017
Earlier work this paper cites.
Ddco: Discovery of deep continuous options for robot learning from demonstrations
Sanjay Krishnan, Roy Fox, Ion Stoica, and Ken Goldberg · 2017
Earlier work this paper cites.
End-to-end driving via conditional imitation learning
Felipe Codevilla, Matthias Müller, Antonio López, Vladlen Koltun, and Alexey Dosovitskiy · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Earlier work this paper cites.
Learning to reach goals via iterated supervised learning
Dibya Ghosh, Abhishek Gupta, Ashwin Reddy, Justin Fu, Coline Devin, Benjamin Eysenbach, and Sergey Levine · 2019
Earlier work this paper cites.
Recurrent independent mechanisms
Anirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani, Sergey Levine, Yoshua Bengio, and Bernhard Schölkopf · 2019
Earlier work this paper cites.
Compile: Compositional imitation learning and execution
Thomas Kipf, Yujia Li, Hanjun Dai, Vinicius Zambaldi, Alvaro Sanchez-Gonzalez, Edward Grefenstette, Pushmeet Kohli, and Peter Battaglia · 2019
Earlier work this paper cites.
Aviral Kumar, Xue Bin Peng, and Sergey Levine · 2019
Earlier work this paper cites.
Reinforcement learning upside down: Don’t predict rewards–just map them to actions
Juergen Schmidhuber · 2019
Earlier work this paper cites.
Training agents using upside-down reinforcement learning
Rupesh Kumar Srivastava, Pranav Shyam, Filipe Mutz, Wojciech Jaśkowski, and Jürgen Schmidhuber · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Cited alongside, same era.
An optimistic perspective on offline reinforcement learning
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi · 2020
Cited alongside, same era.
Opal: Offline primitive discovery for accelerating offline reinforcement learning
Anurag Ajay, Aviral Kumar, Pulkit Agrawal, Sergey Levine, and Ofir Nachum · 2020
Cited alongside, same era.
Hierarchical few-shot imitation with skill transition models
Kourosh Hakhamaneshi, Ruihan Zhao, Albert Zhan, Pieter Abbeel, and Michael Laskin · 2021
Later among the works it cites.
Should i run offline reinforcement learning or behavioral cloning?
Aviral Kumar, Joey Hong, Anikait Singh, and Sergey Levine · 2021
Later among the works it cites.
Policy gradients incorporating the future
David Venuto, Elaine Lau, Doina Precup, and Ofir Nachum · 2021
Later among the works it cites.
Skill preferences: Learning to extract and execute robotic skills from human feedback
Xiaofei Wang, Kimin Lee, Kourosh Hakhamaneshi, Pieter Abbeel, and Michael Laskin · 2021
Later among the works it cites.
Video pretraining (vpt): Learning to act by watching unlabeled online videos
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Learning latent plans from play
Corey Lynch, Mohi Khansari, Ted Xiao, Vikash Kumar, Jonathan Tompson, Sergey Levine, and Pierre Sermanet · 2020
Cited alongside, same era.
Planning from pixels using inverse dynamics models
Keiran Paster, Sheila A McIlraith, and Jimmy Ba · 2020
Cited alongside, same era.
Accelerating reinforcement learning with learned skill priors
Karl Pertsch, Youngwoon Lee, and Joseph J Lim · 2020
Cited alongside, same era.
Learning robot skills with temporal variational inference
Tanmay Shankar and Abhinav Gupta · 2020
Cited alongside, same era.
Parrot: Data-driven behavioral priors for reinforcement learning
Avi Singh, Huihan Liu, Gaoyue Zhou, Albert Yu, Nicholas Rhinehart, and Sergey Levine · 2020
Cited alongside, same era.
Plas: Latent action space for offline reinforcement learning
Wenxuan Zhou, Sujay Bajracharya, and David Held · 2020
Cited alongside, same era.
Bowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, and Jeff Clune · 2022
Closest in time.
When does return-conditioned supervised learning work for offline reinforcement learning?
David Brandfonbrener, Alberto Bietti, Jacob Buckman, Romain Laroche, and Joan Bruna · 2022
Closest in time.
Imitating past successes can be very suboptimal
Benjamin Eysenbach, Soumith Udatha, Sergey Levine, and Ruslan Salakhutdinov · 2022
Closest in time.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
Linxi Fan, Guanzhi Wang, Yunfan Jiang, Ajay Mandlekar, Yuncong Yang, Haoyi Zhu, Andrew Tang, De-An Huang, Yuke Zhu, and Anima Anandkumar · 2022
Closest in time.
Multi-game decision transformers
Kuang-Huei Lee, Ofir Nachum, Mengjiao Yang, Lisa Lee, Daniel Freeman, Winnie Xu, Sergio Guadarrama, Ian Fischer, Eric Jang, Henryk Michalewski, et al · 2022
Closest in time.
You can’t count on luck: Why decision transformers fail in stochastic environments
Keiran Paster, Sheila McIlraith, and Jimmy Ba · 2022
Closest in time.
Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, et al · 2022
Closest in time.
Can wikipedia help offline reinforcement learning?
Machel Reid, Yutaro Yamada, and Shixiang Shane Gu · 2022
Closest in time.
Upside-down reinforcement learning can diverge in stochastic environments with episodic resets
Miroslav Štrupl, Francesco Faccio, Dylan R Ashley, Jürgen Schmidhuber, and Rupesh Kumar Srivastava · 2022
Closest in time.
Addressing optimism bias in sequence modeling for reinforcement learning
Adam R Villaflor, Zhe Huang, Swapnil Pande, John M Dolan, and Jeff Schneider · 2022
Closest in time.