Fetching the paper…
Reading the bibliography…
Diffusion models have achieved remarkable success in sequential decision-making by leveraging the highly expressive model capabilities in policy learning.
Rank analysis of incomplete block designs: I. The method of paired comparisons
Bradley, R. A.; and Terry, M. E. 1952 · 1952
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J.; Kumar, A.; Nachum, O.; Tucker, G.; and Levine, S. 2020 · 2004
Earlier work this paper cites.
Batch reinforcement learning
Lange, S.; Gabel, T.; and Riedmiller, M. 2012 · 2012
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T.; Zhou, A.; Abbeel, P.; and Levine, S. 2018 · 2018
Earlier work this paper cites.
Training deep spiking convolutional neural networks with stdp-based unsupervised pre-training followed by supervised fine-tuning
Lee, C.; Panda, P.; Srinivasan, G.; and Roy, K. 2018 · 2018
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Earlier work this paper cites.
Denoising Diffusion Implicit Models
Song, J.; Meng, C.; and Ermon, S. 2020 · 2020
Earlier work this paper cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T.; Quillen, D.; He, Z.; Julian, R.; Hausman, K.; Finn, C.; and Levine, S. 2020 · 2020
Earlier work this paper cites.
Classifier-Free Diffusion Guidance
Ho, J.; and Salimans, T. 2021 · 2021
Earlier work this paper cites.
B-Pref: Benchmarking Preference-Based Reinforcement Learning
Lee, K.; Smith, L.; Dragan, A.; and Abbeel, P. 2021 · 2021
Earlier work this paper cites.
Improved denoising diffusion probabilistic models
Nichol, A. Q.; and Dhariwal, P. 2021 · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y.; Jones, A.; Ndousse, K.; Askell, A.; Chen, A.; DasSarma, N.; Drain, D.; Fort, S.; Ganguli, D.; Henighan, T.; et al. 2022 · 2022
Earlier work this paper cites.
Planning with Diffusion for Flexible Behavior Synthesis
Janner, M.; Du, Y.; Tenenbaum, J. B.; and Levine, S. 2022 · 2022
Earlier work this paper cites.
Offline Reinforcement Learning with Implicit Q-Learning
Kostrikov, I.; Nair, A.; and Levine, S. 2022 · 2022
Earlier work this paper cites.
Meta-Reward-Net: Implicitly Differentiable Reward Learning for Preference-based Reinforcement Learning
Liu, R.; Bai, F.; Du, Y.; and Yang, Y. 2022 · 2022
Earlier work this paper cites.
Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps
Lu, C.; Zhou, Y.; Bao, F.; Chen, J.; Li, C.; and Zhu, J. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022 · 2022
Earlier work this paper cites.
NeoRL: A Near Real-World Benchmark for Offline Reinforcement Learning
Qin, R.-J.; Zhang, X.; Gao, S.; Chen, X.-H.; Li, Z.; Zhang, W.; and Yu, Y. 2022 · 2022
Earlier work this paper cites.
Is Conditional Generative Modeling all you need for Decision Making?
Ajay, A.; Du, Y.; Gupta, A.; Tenenbaum, J. B.; Jaakkola, T. S.; and Agrawal, P. 2023 · 2023
Earlier work this paper cites.
Direct Preference-based Policy Optimization without Reward Modeling
An, G.; Lee, J.; Zuo, X.; Kosaka, N.; Kim, K.-M.; and Song, H. O. 2023 · 2023
Earlier work this paper cites.
Training Diffusion Models with Reinforcement Learning
Black, K.; Janner, M.; Du, Y.; Kostrikov, I.; and Levine, S. 2023 · 2023
Earlier work this paper cites.
TextDiffuser: Diffusion Models as Text Painters
Chen, J.; Huang, Y.; Lv, T.; Cui, L.; Chen, Q.; and Wei, F. 2023 · 2023
Earlier work this paper cites.
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
Chi, C.; Feng, S.; Du, Y.; Xu, Z.; Cousineau, E.; Burchfiel, B.; and Song, S. 2023 · 2023
Cited alongside, same era.
Directly fine-tuning diffusion models on differentiable rewards
Clark, K.; Vicol, P.; Swersky, K.; and Fleet, D. J. 2023 · 2023
Cited alongside, same era.
Learning Universal Policies via Text-Guided Video Generation
Du, Y.; Yang, S.; Dai, B.; Dai, H.; Nachum, O.; Tenenbaum, J. B.; Schuurmans, D.; and Abbeel, P. 2023 · 2023
Cited alongside, same era.
Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models
Fan, Y.; Watkins, O.; Du, Y.; Liu, H.; Ryu, M.; Boutilier, C.; Abbeel, P.; Ghavamzadeh, M.; Lee, K.; and Lee, K. 2023 · 2023
Cited alongside, same era.
ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills
Gu, J.; Xiang, F.; Li, X.; Ling, Z.; Liu, X.; Mu, T.; Tang, Y.; Tao, S.; Wei, X.; Yao, Y.; Yuan, X.; Xie, P.; Huang, Z.; Chen, R.; and Su, H. 2023 · 2023
Cited alongside, same era.
Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning
Wang, Z.; Hunt, J. J.; and Zhou, M. 2023 · 2023
Later among the works it cites.
Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model
Yang, K.; Tao, J.; Lyu, J.; Ge, C.; Chen, J.; Li, Q.; Shen, W.; Zhu, X.; and Li, X. 2023 · 2023
Later among the works it cites.
Reward-directed conditional diffusion: Provable distribution estimation and reward improvement
Yuan, H.; Huang, K.; Ni, C.; Chen, M.; and Wang, M. 2023 · 2023
Later among the works it cites.
Diffusion models for reinforcement learning: A survey
Zhu, Z.; Zhao, H.; He, H.; Zhong, Y.; Zhang, S.; Yu, Y.; and Zhang, W. 2023 · 2023
Later among the works it cites.
Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Idql: Implicit q-learning as an actor-critic method with diffusion policies
Hansen-Estruch, P.; Kostrikov, I.; Janner, M.; Kuba, J. G.; and Levine, S. 2023 · 2023
Cited alongside, same era.
Diffusion Model is an Effective Planner and Data Synthesizer for Multi-Task Reinforcement Learning
He, H.; Bai, C.; Xu, K.; Yang, Z.; Zhang, W.; Wang, D.; Zhao, B.; and Li, X. 2023 · 2023
Cited alongside, same era.
Contrastive Prefence Learning: Learning from Human Feedback without RL
Hejna, J.; Rafailov, R.; Sikchi, H.; Finn, C.; Niekum, S.; Knox, W. B.; and Sadigh, D. 2023 · 2023
Cited alongside, same era.
Efficient diffusion policies for offline reinforcement learning
Kang, B.; Ma, X.; Du, C.; Pang, T.; and Yan, S. 2023 · 2023
Cited alongside, same era.
Imagic: Text-based real image editing with diffusion models
Kawar, B.; Zada, S.; Lang, O.; Tov, O.; Chang, H.; Dekel, T.; Mosseri, I.; and Irani, M. 2023 · 2023
Cited alongside, same era.
Preference Transformer: Modeling Human Preferences using Transformers for RL
Kim, C.; Park, J.; Shin, J.; Lee, H.; Abbeel, P.; and Lee, K. 2023 · 2023
Cited alongside, same era.
ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Li, Z.; Xu, T.; Zhang, Y.; Yu, Y.; Sun, R.; and Luo, Z.-Q. 2023 · 2023
Cited alongside, same era.
Ahmadian, A.; Cremer, C.; Gallé, M.; Fadaee, M.; Kreutzer, J.; Üstün, A.; and Hooker, S. 2024 · 2024
Closest in time.
Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF
Cen, S.; Mei, J.; Goshvadi, K.; Dai, H.; Yang, T.; Yang, S.; Schuurmans, D.; Chi, Y.; and Dai, B. 2024 · 2024
Closest in time.
AlignDiff: Aligning Diverse Human Preferences via Behavior-Customisable Diffusion Model
Dong, Z.; Yuan, Y.; HAO, J.; Ni, F.; Mu, Y.; ZHENG, Y.; Hu, Y.; Lv, T.; Fan, C.; and Hu, Z. 2024 · 2024
Closest in time.
Video Language Planning
Du, Y.; Yang, S.; Florence, P.; Xia, F.; Wahid, A.; brian ichter; Sermanet, P.; Yu, T.; Abbeel, P.; Tenenbaum, J. B.; Kaelbling, L. P.; Zeng, A.; and Tompson, J. 2024 · 2024
Closest in time.
Task-agnostic Pre-training and Task-guided Fine-tuning for Versatile Diffusion Planner
Fan, C.; Bai, C.; Shan, Z.; He, H.; Zhang, Y.; and Wang, Z. 2024 · 2024
Closest in time.
Learning an actionable discrete diffusion policy via large-scale actionless video pre-training
He, H.; Bai, C.; Pan, L.; Zhang, W.; Zhao, B.; and Li, X. 2024 · 2024
Closest in time.
Inverse preference learning: Preference-based rl without a reward function
Hejna, J.; and Sadigh, D. 2024 · 2024
Closest in time.
Huang, A.; Zhan, W.; Xie, T.; Lee, J. D.; Sun, W.; Krishnamurthy, A.; and Foster, D. J. 2024 · 2024
Closest in time.
Models of human preference for learning reward functions
Knox, W. B.; Hatgis-Kessell, S.; Booth, S.; Niekum, S.; Stone, P.; and Allievi, A. G. 2024 · 2024
Closest in time.
Rewardbench: Evaluating reward models for language modeling
Lambert, N.; Pyatkin, V.; Morrison, J.; Miranda, L.; Lin, B. Y.; Chandu, K.; Dziri, N.; Kumar, S.; Zick, T.; Choi, Y.; et al. 2024 · 2024
Closest in time.
Skilldiffuser: Interpretable hierarchical planning via skill abstractions in diffusion-based task execution
Liang, Z.; Mu, Y.; Ma, H.; Tomizuka, M.; Ding, M.; and Luo, P. 2024 · 2024
Closest in time.
Inverse-RLignment: Inverse Reinforcement Learning from Demonstrations for LLM Alignment
Sun, H.; and van der Schaar, M. 2024 · 2024
Closest in time.
Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Wang, C.; Jiang, Y.; Yang, C.; Liu, H.; and Chen, Y. 2024 · 2024
Closest in time.
Learning Interactive Real-World Simulators
Yang, S.; Du, Y.; Ghasemipour, S. K. S.; Tompson, J.; Kaelbling, L. P.; Schuurmans, D.; and Abbeel, P. 2024 · 2024
Closest in time.
Regularized Conditional Diffusion Model for Multi-Task Preference Alignment
Yu, X.; Bai, C.; He, H.; Wang, C.; and Li, X. 2024 · 2024
Closest in time.
3d diffusion policy
Ze, Y.; Zhang, G.; Zhang, K.; Hu, C.; Wang, M.; and Xu, H. 2024 · 2024
Closest in time.
Off-policy deep reinforcement learning without exploration
Fujimoto, S.; Meger, D.; and Precup, D. 2019 · 2062
Closest in time.