Fetching the paper…
Reading the bibliography…
Offline goal-conditioned reinforcement learning (GCRL) promises general-purpose skill learning in the form of reaching diverse goals from purely offline datasets.
Convex analysis
R. Tyrrell Rockafellar · 1970
Earlier work this paper cites.
Convex optimization
Stephen Boyd, Stephen P Boyd, and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, Anind K Dey, et al · 2008
Earlier work this paper cites.
Learning bounds for importance weighting
Corinna Cortes, Yishay Mansour, and Mehryar Mohri · 2010
Earlier work this paper cites.
Relative entropy policy search
Jan Peters, Katharina Mulling, and Yasemin Altun · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Batch reinforcement learning
Sascha Lange, Thomas Gabel, and Martin Riedmiller · 2012
Earlier work this paper cites.
Generative adversarial networks, 2014
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Earlier work this paper cites.
Learning from conditional distributions via dual embeddings, 2016
Bo Dai, Niao He, Yunpeng Pan, Byron Boots, and Le Song · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Earlier work this paper cites.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba · 2017
Earlier work this paper cites.
Improved training of wasserstein gans
Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron C Courville · 2017
Earlier work this paper cites.
Multi-goal reinforcement learning: Challenging robotics environments and request for research
Matthias Plappert, Marcin Andrychowicz, Alex Ray, Bob McGrew, Bowen Baker, Glenn Powell, Jonas Schneider, Josh Tobin, Maciek Chociej, Peter Welinder, et al · 2018
Earlier work this paper cites.
Generative adversarial imitation from observation
Faraz Torabi, Garrett Warnell, and Peter Stone · 2018
Cited alongside, same era.
Reinforcement learning: Theory and algorithms
Alekh Agarwal, Nan Jiang, and Sham M Kakade · 2019
Cited alongside, same era.
Robonet: Large-scale multi-robot learning
Sudeep Dasari, Frederik Ebert, Stephen Tian, Suraj Nair, Bernadette Bucher, Karl Schmeckpeper, Siddharth Singh, Sergey Levine, and Chelsea Finn · 2019
Cited alongside, same era.
A divergence minimization perspective on imitation learning methods, 2019
Seyed Kamyar Seyed Ghasemipour, Richard Zemel, and Shixiang Gu · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, George Tucker, and Sergey Levine · 2019
Learning latent plans from play
Corey Lynch, Mohi Khansari, Ted Xiao, Vikash Kumar, Jonathan Tompson, Sergey Levine, and Pierre Sermanet · 2020
Later among the works it cites.
Reinforcement learning via fenchel-rockafellar duality, 2020
Ofir Nachum and Bo Dai · 2020
Later among the works it cites.
Error bounds of imitating policies and environments
Tian Xu, Ziniu Li, and Yang Yu · 2020
Later among the works it cites.
Off-policy imitation learning from observations
Zhuangdi Zhu, Kaixiang Lin, Bo Dai, and Jiayu Zhou · 2020
Later among the works it cites.
Goal-conditioned reinforcement learning with imagined subgoals
Elliot Chane-Sane, Cordelia Schmid, and Ivan Laptev · 2021
Later among the works it cites.
Actionable models: Unsupervised offline reinforcement learning of robotic skills
Yevgen Chebotar, Karol Hausman, Yao Lu, Ted Xiao, Dmitry Kalashnikov, Jake Varley, Alex Irpan, Benjamin Eysenbach, Ryan Julian, Chelsea Finn, et al · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Efficient exploration via state marginal matching
Lisa Lee, Benjamin Eysenbach, Emilio Parisotto, Eric Xing, Sergey Levine, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Ofir Nachum, Yinlam Chow, Bo Dai, and Lihong Li · 2019
Cited alongside, same era.
Algaedice: Policy gradient from arbitrary experience
Ofir Nachum, Bo Dai, Ilya Kostrikov, Yinlam Chow, Lihong Li, and Dale Schuurmans · 2019
Cited alongside, same era.
Planning with goal-conditioned policies
Soroush Nasiriany, Vitchyr Pong, Steven Lin, and Sergey Levine · 2019
Cited alongside, same era.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2019
Cited alongside, same era.
Robel: Robotics benchmarks for learning with low-cost robots
Michael Ahn, Henry Zhu, Kristian Hartikainen, Hugo Ponte, Abhishek Gupta, Sergey Levine, and Vikash Kumar · 2020
Cited alongside, same era.
Rewriting history with inverse rl: Hindsight inference for policy improvement
Ben Eysenbach, Xinyang Geng, Sergey Levine, and Russ R Salakhutdinov · 2020
Cited alongside, same era.
Later among the works it cites.
Multi-goal reinforcement learning environments for simulated franka emika panda robot
Quentin Gallouédec, Nicolas Cazin, Emmanuel Dellandréa, and Liming Chen · 2021
Later among the works it cites.
Learning to reach goals via iterated supervised learning
Dibya Ghosh, Abhishek Gupta, Ashwin Reddy, Justin Fu, Coline Manon Devin, Benjamin Eysenbach, and Sergey Levine · 2021
Later among the works it cites.
Mt-opt: Continuous multi-task robotic reinforcement learning at scale
Dmitry Kalashnikov, Jacob Varley, Yevgen Chebotar, Benjamin Swanson, Rico Jonschkowski, Chelsea Finn, Sergey Levine, and Karol Hausman · 2021
Later among the works it cites.
Optidice: Offline policy optimization via stationary distribution correction estimation
Jongmin Lee, Wonseok Jeon, Byung-Jun Lee, Joelle Pineau, and Kee-Eung Kim · 2021
Later among the works it cites.
Conservative offline distributional reinforcement learning
Yecheng Ma, Dinesh Jayaraman, and Osbert Bastani · 2021
Later among the works it cites.
Universal value density estimation for imitation learning and goal-conditioned reinforcement learning, 2021
Yannick Schroecker and Charles Lee Isbell · 2021
Later among the works it cites.
Towards hyperparameter-free policy selection for offline reinforcement learning
Siyuan Zhang and Nan Jiang · 2021
Later among the works it cites.
DemoDICE: Offline imitation learning with supplementary imperfect demonstrations
Geon-Hyeong Kim, Seokin Seo, Jongmin Lee, Wonseok Jeon, HyeongJoo Hwang, Hongseok Yang, and Kee-Eung Kim · 2022
Closest in time.
Smodice: Versatile offline imitation learning via state occupancy matching
Yecheng Jason Ma, Andrew Shen, Dinesh Jayaraman, and Osbert Bastani · 2022
Closest in time.
Rethinking goal-conditioned supervised learning and its connection to offline RL
Rui Yang, Yiming Lu, Wenzhe Li, Hao Sun, Meng Fang, Yali Du, Xiu Li, Lei Han, and Chongjie Zhang · 2022
Closest in time.