Fetching the paper…
Reading the bibliography…
Due to its training stability and strong expression, the diffusion model has attracted considerable attention in offline reinforcement learning.
Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine. 2019 · 1910
Earlier work this paper cites.
Behavior Regularized Offline Reinforcement Learning
Yifan Wu, George Tucker, and Ofir Nachum. 2019 · 1911
Earlier work this paper cites.
AlgaeDICE: Policy Gradient from Arbitrary Experience
Ofir Nachum, Bo Dai, Ilya Kostrikov, Yinlam Chow, Lihong Li, and Dale Schuurmans. 2019 · 1912
Earlier work this paper cites.
D4RL: Datasets for Deep Data-Driven Reinforcement Learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine. 2020 · 2004
Earlier work this paper cites.
Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu. 2020 · 2005
Earlier work this paper cites.
Relative Entropy Policy Search. In Proceedings of the Twenty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2010, Atlanta, Georgia, USA, July 11-15, 2010 , Maria Fox and David Poole (Eds.). AAAI Press
Jan Peters, Katharina Mülling, and Yasemin Altun. 2010 · 2010
Earlier work this paper cites.
Double Q-learning. In Advances in Neural Information Processing Systems 23: 24th Annual Conference on Neural Information Processing Systems 2010. Proceedings of a meeting held 6-9 December 2010, Vancouver, British Columbia, Canada , John D. Lafferty, Christopher K. I. Williams, John Shawe-Taylor, Richard S. Zemel, and Aron Culotta (Eds.). Curran Associates, Inc., 2613–2621
Hado van Hasselt. 2010 · 2010
Earlier work this paper cites.
Batch Reinforcement Learning
Sascha Lange, Thomas Gabel, and Martin A. Riedmiller. 2012 · 2012
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control. In 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS 2012, Vilamoura, Algarve, Portugal, October 7-12, 2012 . IEEE, 5026–5033
Emanuel Todorov, Tom Erez, and Yuval Tassa. 2012 · 2012
Earlier work this paper cites.
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2016 · 2016
Earlier work this paper cites.
A Distributional Perspective on Reinforcement Learning. In Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017 (Proceedings of Machine Learning Research, Vol. 70) , Doina Precup and Yee Whye Teh (Eds.). PMLR, 449–458
Marc G. Bellemare, Will Dabney, and Rémi Munos. 2017 · 2017
Earlier work this paper cites.
Proximal Policy Optimization Algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Earlier work this paper cites.
Maximum a Posteriori Policy Optimisation. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings . OpenReview.net
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Rémi Munos, Nicolas Heess, and Martin A. Riedmiller. 2018 · 2018
Earlier work this paper cites.
Distributed Distributional Deterministic Policy Gradients. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings . OpenReview.net
Gabriel Barth-Maron, Matthew W. Hoffman, David Budden, Will Dabney, Dan Horgan, Dhruva TB, Alistair Muldal, Nicolas Heess, and Timothy P. Lillicrap. 2018 · 2018
Earlier work this paper cites.
Addressing Function Approximation Error in Actor-Critic Methods. In Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10-15, 2018 (Proceedings of Machine Learning Research, Vol. 80) , Jennifer G. Dy and Andreas Krause (Eds.). PMLR, 1582–1591
Scott Fujimoto, Herke van Hoof, and David Meger. 2018 · 2018
Earlier work this paper cites.
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. In Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10-15, 2018 (Proceedings of Machine Learning Research, Vol. 80) , Jennifer G. Dy and Andreas Krause (Eds.). PMLR, 1856–1865
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018 · 2018
Earlier work this paper cites.
Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada , Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’Alché-Buc, Emily B. Fox, and Roman Garnett (Eds.). 11761–11771
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine. 2019 · 2019
Earlier work this paper cites.
A distributional view on multi-objective policy optimization. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event (Proceedings of Machine Learning Research, Vol. 119) . PMLR, 11–22
Abbas Abdolmaleki, Sandy H. Huang, Leonard Hasenclever, Michael Neunert, H. Francis Song, Martina Zambelli, Murilo F. Martins, Nicolas Heess, Raia Hadsell, and Martin A. Riedmiller. 2020 · 2020
Earlier work this paper cites.
Denoising Diffusion Probabilistic Models. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual , Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (Eds.)
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Cited alongside, same era.
Conservative Q-Learning for Offline Reinforcement Learning. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual , Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (Eds.)
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. 2020 · 2020
Cited alongside, same era.
dm_control: Software and tasks for continuous control
Saran Tunyasuvunakool, Alistair Muldal, Yotam Doron, Siqi Liu, Steven Bohez, Josh Merel, Tom Erez, Timothy P. Lillicrap, Nicolas Heess, and Yuval Tassa. 2020 · 2020
Cited alongside, same era.
Mastering Diverse Domains through World Models
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy P. Lillicrap. 2023 · 2023
Closest in time.
IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies
Philippe Hansen-Estruch, Ilya Kostrikov, Michael Janner, Jakub Grudzien Kuba, and Sergey Levine. 2023 · 2023
Closest in time.
Efficient Diffusion Policies for Offline Reinforcement Learning
Bingyi Kang, Xiao Ma, Chao Du, Tianyu Pang, and Shuicheng Yan. 2023 · 2023
Closest in time.
DALL-E-Bot: Introducing Web-Scale Diffusion Models to Robotics
Ivan Kapelyukh, Vitalis Vosylius, and Edward Johns. 2023 · 2023
Closest in time.
Crossway Diffusion: Improving Diffusion-based Visuomotor Policy via Self-supervised Learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Diffusion Models Beat GANs on Image Synthesis. In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual , Marc’Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan (Eds.). 8780–8794
Prafulla Dhariwal and Alexander Quinn Nichol. 2021 · 2021
Cited alongside, same era.
A Minimalist Approach to Offline Reinforcement Learning. In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual , Marc’Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan (Eds.). 20132–20145
Scott Fujimoto and Shixiang Shane Gu. 2021 · 2021
Cited alongside, same era.
Score-Based Generative Modeling through Stochastic Differential Equations. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenReview.net
Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. 2021 · 2021
Cited alongside, same era.
Offline Reinforcement Learning for Autonomous Driving with Real World Driving Data. In 25th IEEE International Conference on Intelligent Transportation Systems, ITSC 2022, Macau, China, October 8-12, 2022 . IEEE, 3417–3422
Xing Fang, Qichao Zhang, Yinfeng Gao, and Dongbin Zhao. 2022 · 2022
Cited alongside, same era.
Know Your Boundaries: The Necessity of Explicit Behavioral Cloning in Offline RL
Wonjoon Goo and Scott Niekum. 2022 · 2022
Cited alongside, same era.
Planning with Diffusion for Flexible Behavior Synthesis. In International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA (Proceedings of Machine Learning Research, Vol. 162) , Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvári, Gang Niu, and Sivan Sabato (Eds.). PMLR, 9902–9915
Michael Janner, Yilun Du, Joshua B. Tenenbaum, and Sergey Levine. 2022 · 2022
Cited alongside, same era.
Elucidating the Design Space of Diffusion-Based Generative Models. In NeurIPS
Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. 2022 · 2022
Cited alongside, same era.
Offline Reinforcement Learning with Implicit Q-Learning. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022 . OpenReview.net
Ilya Kostrikov, Ashvin Nair, and Sergey Levine. 2022 · 2022
Cited alongside, same era.
Offline Reinforcement Learning with Value-based Episodic Memory. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022 . OpenReview.net
Xiaoteng Ma, Yiqin Yang, Hao Hu, Jun Yang, Chongjie Zhang, Qianchuan Zhao, Bin Liang, and Qihan Liu. 2022 · 2022
Cited alongside, same era.
Xiang Li, Varun Belagali, Jinghuan Shang, and Michael S. Ryoo. 2023 · 2023
Closest in time.
Cong Lu, Philip J. Ball, and Jack Parker-Holder. 2023a · 2023
Closest in time.
Cheng Lu, Huayu Chen, Jianfei Chen, Hang Su, Chongxuan Li, and Jun Zhu. 2023b · 2023
Closest in time.
Imitating Human Behaviour with Diffusion Models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023 . OpenReview.net
Tim Pearce, Tabish Rashid, Anssi Kanervisto, David Bignell, Mingfei Sun, Raluca Georgescu, Sergio Valcarcel Macua, Shan Zheng Tan, Ida Momennejad, Katja Hofmann, and Sam Devlin. 2023 · 2023
Closest in time.
Trace and Pace: Controllable Pedestrian Animation via Guided Trajectory Diffusion. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023 . IEEE, 13756–13766
Davis Rempe, Zhengyi Luo, Xue Bin Peng, Ye Yuan, Kris Kitani, Karsten Kreis, Sanja Fidler, and Or Litany. 2023 · 2023
Closest in time.
Goal-Conditioned Imitation Learning using Score-based Diffusion Policies. In Robotics: Science and Systems XIX, Daegu, Republic of Korea, July 10-14, 2023 , Kostas E. Bekris, Kris Hauser, Sylvia L. Herbert, and Jingjin Yu (Eds.)
Moritz Reuss, Maximilian Li, Xiaogang Jia, and Rudolf Lioutikov. 2023 · 2023
Closest in time.
Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever. 2023 · 2023
Closest in time.
Diffusion Model-Augmented Behavioral Cloning
Hsiang-Chun Wang, Shang-Fu Chen, and Shao-Hua Sun. 2023a · 2023
Closest in time.
Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023 . OpenReview.net
Zhendong Wang, Jonathan J. Hunt, and Mingyuan Zhou. 2023b · 2023
Closest in time.
The In-Sample Softmax for Offline Reinforcement Learning. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023 . OpenReview.net
Chenjun Xiao, Han Wang, Yangchen Pan, Adam White, and Martha White. 2023 · 2023
Closest in time.
Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023 . OpenReview.net
Haoran Xu, Li Jiang, Jianxiong Li, Zhuoran Yang, Zhaoran Wang, Wai Kin Victor Chan, and Xianyuan Zhan. 2023 · 2023
Closest in time.
Policy Representation via Diffusion Probability Model for Reinforcement Learning
Long Yang, Zhixiong Huang, Fenghao Lei, Yucun Zhong, Yiming Yang, Cong Fang, Shiting Wen, Binbin Zhou, and Zhouchen Lin. 2023 · 2023
Closest in time.
Ziyuan Zhong, Davis Rempe, Danfei Xu, Yuxiao Chen, Sushant Veer, Tong Che, Baishakhi Ray, and Marco Pavone. 2023 · 2023
Closest in time.
Off-Policy Deep Reinforcement Learning without Exploration. In Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA (Proceedings of Machine Learning Research, Vol. 97) , Kamalika Chaudhuri and Ruslan Salakhutdinov (Eds.). PMLR, 2052–2062
Scott Fujimoto, David Meger, and Doina Precup. 2019 · 2062
Closest in time.