Fetching the paper…
Reading the bibliography…
Generative models such as diffusion have been employed as world models in offline reinforcement learning to generate synthetic data for more effective learning.
David Ha and Jürgen Schmidhuber · 2018
Earlier work this paper cites.
Overcoming exploration in reinforcement learning with demonstrations
Ashvin Nair, Bob McGrew, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine · 2019
Earlier work this paper cites.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Earlier work this paper cites.
Model-based reinforcement learning for atari
Lukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski, Roy H Campbell, Konrad Czechowski, Dumitru Erhan, Chelsea Finn, Piotr Kozakowski, Sergey Levine, et al · 2019
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi · 2019
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Earlier work this paper cites.
Deployment-efficient reinforcement learning via model-based offline optimization
Tatsuya Matsushima, Hiroki Furuta, Yutaka Matsuo, Ofir Nachum, and Shixiang Gu · 2020
Earlier work this paper cites.
Mopo: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Y Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Earlier work this paper cites.
Morel: Model-based offline reinforcement learning
Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims · 2020
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Earlier work this paper cites.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Earlier work this paper cites.
Regularizing model-based planning with energy-based models
Rinu Boney, Juho Kannala, and Alexander Ilin · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Deep reinforcement learning for autonomous driving: A survey
B Ravi Kiran, Ibrahim Sobh, Victor Talpaert, Patrick Mannion, Ahmad A Al Sallab, Senthil Yogamani, and Patrick Pérez · 2021
Earlier work this paper cites.
Combo: Conservative offline model-based policy optimization
Tianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran, Sergey Levine, and Chelsea Finn · 2021
Earlier work this paper cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Paria Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao, and Stuart Russell · 2021
Cited alongside, same era.
Vector quantized models for planning
Sherjil Ozair, Yazhe Li, Ali Razavi, Ioannis Antonoglou, Aaron Van Den Oord, and Oriol Vinyals · 2021
Cited alongside, same era.
Offline reinforcement learning as one big sequence modeling problem
Michael Janner, Qiyang Li, and Sergey Levine · 2021
Cited alongside, same era.
Offline reinforcement learning with implicit q-learning
Ilya Kostrikov, Ashvin Nair, and Sergey Levine · 2021
Cited alongside, same era.
A minimalist approach to offline reinforcement learning
Scott Fujimoto and Shixiang Shane Gu · 2021
Cited alongside, same era.
Improved denoising diffusion probabilistic models
Cong Lu, Philip J Ball, and Jack Parker-Holder · 2023
Later among the works it cites.
Metadiffuser: Diffusion model as conditional planner for offline meta-rl
Fei Ni, Jianye Hao, Yao Mu, Yifu Yuan, Yan Zheng, Bin Wang, and Zhixuan Liang · 2023
Later among the works it cites.
Hierarchical diffusion for offline decision making
Wenhao Li, Xiangfeng Wang, Bo Jin, and Hongyuan Zha · 2023
Later among the works it cites.
Diffusion model is an effective planner and data synthesizer for multi-task reinforcement learning
Haoran He, Chenjia Bai, Kang Xu, Zhuoran Yang, Weinan Zhang, Dong Wang, Bin Zhao, and Xuelong Li · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexander Quinn Nichol and Prafulla Dhariwal · 2021
Cited alongside, same era.
Bringing fairness to actor-critic reinforcement learning for network utility optimization
Jingdi Chen, Yimeng Wang, and Tian Lan · 2021
Cited alongside, same era.
Rambo-rl: Robust adversarial model-based offline reinforcement learning
Marc Rigter, Bruno Lacerda, and Nick Hawes · 2022
Cited alongside, same era.
Mismatched no more: Joint model-policy optimization for model-based rl
Benjamin Eysenbach, Alexander Khazatsky, Sergey Levine, and Russ R Salakhutdinov · 2022
Cited alongside, same era.
Planning with diffusion for flexible behavior synthesis
Michael Janner, Yilun Du, Joshua B Tenenbaum, and Sergey Levine · 2022
Cited alongside, same era.
Diffusion policies as an expressive policy class for offline reinforcement learning
Zhendong Wang, Jonathan J Hunt, and Mingyuan Zhou · 2022
Cited alongside, same era.
Offline reinforcement learning via high-fidelity generative behavior modeling
Huayu Chen, Cheng Lu, Chengyang Ying, Hang Su, and Jun Zhu · 2022
Cited alongside, same era.
Zhengbang Zhu, Minghuan Liu, Liyuan Mao, Bingyi Kang, Minkai Xu, Yong Yu, Stefano Ermon, and Weinan Zhang · 2023
Later among the works it cites.
Safediffuser: Safe planning with diffusion probabilistic models
Wei Xiao, Tsun-Hsuan Wang, Chuang Gan, and Daniela Rus · 2023
Later among the works it cites.
Extracting reward functions from diffusion models
Felipe Nuti, Tim Franzmeyer, and João F Henriques · 2023
Later among the works it cites.
Live in the moment: Learning dynamics model adapted to evolving policy
Xiyao Wang, Wichayaporn Wongkamjan, Ruonan Jia, and Furong Huang · 2023
Later among the works it cites.
Osteoporotic-like vertebral fracture with less than 20% height loss is associated with increased further vertebral fracture risk in older women: the mros and msos (hong kong) year-18 follow-up radiograph results
Yì Xiáng J Wáng, Zhi-Hui Lu, Jason CS Leung, Ze-Yu Fang, and Timothy CY Kwok · 2023
Later among the works it cites.
Implementing first-person shooter game ai in wild-scav with rule-enhanced deep reinforcement learning
Zeyu Fang, Jian Zhao, Wengang Zhou, and Houqiang Li · 2023
Later among the works it cites.
Value functions factorization with latent state information sharing in decentralized multi-agent policy gradients
Hanhan Zhou, Tian Lan, and Vaneet Aggarwal · 2023
Later among the works it cites.
Mac-po: Multi-agent experience replay via collective priority optimization
Yongsheng Mei, Hanhan Zhou, Tian Lan, Guru Venkataramani, and Peng Wei · 2023
Later among the works it cites.
Real-time network intrusion detection via decision transformers
Jingdi Chen, Hanhan Zhou, Yongsheng Mei, Gina Adam, Nathaniel D Bastian, and Tian Lan · 2023
Later among the works it cites.
Zihan Ding, Amy Zhang, Yuandong Tian, and Qinqing Zheng · 2024
Closest in time.
Generating behaviorally diverse policies with latent diffusion models
Shashank Hegde, Sumeet Batra, KR Zentner, and Gaurav Sukhatme · 2024
Closest in time.
A distributed abstract mac layer for cooperative learning on internet of vehicles
Yifei Zou, Zuyuan Zhang, Congwei Zhang, Yanwei Zheng, Dongxiao Yu, and Jiguo Yu · 2024
Closest in time.
Cooperative backdoor attack in decentralized reinforcement learning with theoretical guarantee
Mengtong Gao, Yifei Zou, Zuyuan Zhang, Xiuzhen Cheng, and Dongxiao Yu · 2024
Closest in time.