Fetching the paper…
Reading the bibliography…
We introduce Diffusion Policy Policy Optimization, DPPO, an algorithmic framework including best practices for fine-tuning diffusion-based policies (e.g.
Alvinn: An autonomous land vehicle in a neural network
D. A. Pomerleau · 1988
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour · 1999
Earlier work this paper cites.
Pattern recognition and machine learning
C. M. Bishop and N. M. Nasrabadi · 2006
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
J. Peters and S. Schaal · 2007
Earlier work this paper cites.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2010
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli · 2015
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Testing the manifold hypothesis
C. Fefferman, S. Mitter, and H. Narayanan · 2016
Earlier work this paper cites.
Apriltag 2: Efficient and robust fiducial detection
J. Wang and E. Olson · 2016
Earlier work this paper cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
A. Rajeswaran, V. Kumar, A. Gupta, G. Vezzani, J. Schulman, E. Todorov, and S. Levine · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
M. Vecerik, T. Hester, J. Scholz, F. Wang, O. Pietquin, B. Piot, N. Heess, T. Rothörl, T. Lampe, and M. Riedmiller · 2017
Earlier work this paper cites.
Spinning Up in Deep Reinforcement Learning
J. Achiam · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Deep q-learning from demonstrations
T. Hester, M. Vecerik, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, D. Horgan, J. Quan, A. Sendonaris, I. Osband, et al · 2018
Earlier work this paper cites.
An algorithmic perspective on imitation learning
T. Osa, J. Pajarinen, G. Neumann, J. A. Bagnell, P. Abbeel, J. Peters, et al · 2018
Earlier work this paper cites.
Deep reinforcement learning for de novo drug design
M. Popova, O. Isayev, and A. Tropsha · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
Implementation matters in deep rl: A case study on ppo and trpo
L. Engstrom, A. Ilyas, S. Santurkar, D. Tsipras, F. Janoos, L. Rudolph, and A. Madry · 2019
Earlier work this paper cites.
Self-supervised correspondence in visuomotor policy learning
P. Florence, L. Manuelli, and R. Tedrake · 2019
Earlier work this paper cites.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
A. Gupta, V. Kumar, C. Lynch, S. Levine, and K. Hausman · 2019
Earlier work this paper cites.
Learning agile and dynamic motor skills for legged robots
J. Hwangbo, J. Lee, A. Dosovitskiy, D. Bellicoso, V. Tsounis, V. Koltun, and M. Hutter · 2019
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
X. B. Peng, A. Kumar, G. Zhang, and S. Levine · 2019
Earlier work this paper cites.
Dexterous manipulation with deep reinforcement learning: Efficient, general, and low-cost
H. Zhu, A. Gupta, A. Rajeswaran, S. Levine, and V. Kumar · 2019
Earlier work this paper cites.
Learning dexterous in-hand manipulation
O. M. Andrychowicz, B. Baker, M. Chociej, R. Jozefowicz, B. McGrew, J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, et al · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Cited alongside, same era.
Diffwave: A versatile diffusion model for audio synthesis
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro · 2020
Cited alongside, same era.
Learning active task-oriented exploration policies for bridging the sim-to-real gap
J. Liang, S. Saxena, and O. Kroemer · 2020
Cited alongside, same era.
Awac: Accelerating online reinforcement learning with offline datasets
A. Nair, A. Gupta, M. Dalal, and S. Levine · 2020
Cited alongside, same era.
Residual reinforcement learning from demonstrations
M. Alakuijala, G. Dulac-Arnold, J. Mairal, J. Ponce, and C. Schmid · 2021
K. Lei, Z. He, C. Lu, K. Hu, Y. Gao, and H. Xu · 2023
Later among the works it cites.
Adaptdiffuser: Diffusion models as adaptive self-evolving planners
Z. Liang, Y. Mu, M. Ding, F. Ni, M. Tomizuka, and P. Luo · 2023
Later among the works it cites.
Imitating human behaviour with diffusion models
T. Pearce, T. Rashid, A. Kanervisto, D. Bignell, M. Sun, R. Georgescu, S. V. Macua, S. Z. Tan, I. Momennejad, K. Hofmann, et al · 2023
Later among the works it cites.
Interpreting and improving diffusion models from an optimization perspective
F. Permenter and C. Yuan · 2023
Later among the works it cites.
Learning a diffusion model policy from rewards via q-score matching
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Polymetis
Y. Lin, A. S. Wang, G. Sutanto, A. Rai, and F. Meier · 2021
Cited alongside, same era.
Isaac gym: High performance gpu-based physics simulation for robot learning
V. Makoviychuk, L. Wawrzyniak, Y. Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa, et al · 2021
Cited alongside, same era.
What matters in learning from offline human demonstrations for robot manipulation
A. Mandlekar, D. Xu, J. Wong, S. Nasiriany, C. Wang, R. Kulkarni, L. Fei-Fei, S. Savarese, Y. Zhu, and R. Martín-Martín · 2021
Cited alongside, same era.
Improved denoising diffusion probabilistic models
A. Q. Nichol and P. Dhariwal · 2021
Cited alongside, same era.
Amp: Adversarial motion priors for stylized physics-based character control
X. B. Peng, Z. Ma, P. Abbeel, S. Levine, and A. Kanazawa · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Implicit behavioral cloning
P. Florence, C. Lynch, A. Zeng, O. A. Ramirez, A. Wahid, L. Downs, A. Wong, J. Lee, I. Mordatch, and J. Tompson · 2022
Cited alongside, same era.
M. Psenka, A. Escontrela, P. Abbeel, and Y. Ma · 2023
Later among the works it cites.
AdaptSim: Task-driven simulation adaptation for sim-to-real transfer
A. Z. Ren, H. Dai, B. Burchfiel, and A. Majumdar · 2023
Later among the works it cites.
Goal-conditioned imitation learning using score-based diffusion policies
M. Reuss, M. Li, X. Jia, and R. Lioutikov · 2023
Later among the works it cites.
World models via policy-guided trajectory diffusion
M. Rigter, J. Yamada, and I. Posner · 2023
Later among the works it cites.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
N. Ruiz, Y. Li, V. Jampani, Y. Pritch, M. Rubinstein, and K. Aberman · 2023
Later among the works it cites.
Nomad: Goal masked diffusion policies for navigation and exploration
A. Sridhar, D. Shah, C. Glossop, and S. Levine · 2023
Later among the works it cites.
Reasoning with latent diffusion in offline reinforcement learning
S. Venkatraman, S. Khaitan, R. T. Akella, J. Dolan, J. Schneider, and G. Berseth · 2023
Later among the works it cites.
Diffusion model alignment using direct preference optimization
B. Wallace, M. Dang, R. Rafailov, L. Zhou, A. Lou, S. Purushwalkam, S. Ermon, C. Xiong, S. Joty, and N. Naik · 2023
Later among the works it cites.
Learning fine-grained bimanual manipulation with low-cost hardware
T. Z. Zhao, V. Kumar, S. Levine, and C. Finn · 2023
Later among the works it cites.
Diffusion models for reinforcement learning: A survey
Z. Zhu, H. Zhao, H. He, Y. Zhong, S. Zhang, Y. Yu, and W. Zhang · 2023
Later among the works it cites.
Provable guarantees for generative behavior cloning: Bridging low-level stability and high-level behavior
A. Block, A. Jadbabaie, D. Pfrommer, M. Simchowitz, and R. Tedrake · 2024
Closest in time.
Genie: Generative interactive environments
J. Bruce, M. D. Dennis, A. Edwards, J. Parker-Holder, Y. Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps, et al · 2024
Closest in time.
Tutorial on diffusion models for imaging and vision
S. H. Chan · 2024
Closest in time.
Diffusion forcing: Next-token prediction meets full-sequence diffusion
B. Chen, D. M. Monso, Y. Du, M. Simchowitz, R. Tedrake, and V. Sitzmann · 2024
Closest in time.
Reinforcement learning for fine-tuning text-to-image diffusion models
Y. Fan, O. Watkins, Y. Du, H. Liu, M. Ryu, C. Boutilier, P. Abbeel, M. Ghavamzadeh, K. Lee, and K. Lee · 2024
Closest in time.
Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation
Z. Fu, T. Z. Zhao, and C. Finn · 2024
Closest in time.
Protein-ligand interaction prior for binding-aware 3d molecule diffusion models
Z. Huang, L. Yang, X. Zhou, Z. Zhang, W. Zhang, X. Zheng, J. Chen, Y. Wang, C. Bin, and W. Yang · 2024
Closest in time.
M. T. Jackson, M. T. Matthews, C. Lu, B. Ellis, S. Whiteson, and J. Foerster · 2024
Closest in time.
Towards diverse behaviors: A benchmark for imitation learning with human demonstrations
X. Jia, D. Blessing, X. Jiang, M. Reuss, A. Donat, R. Lioutikov, and G. Neumann · 2024
Closest in time.
Efficient diffusion policies for offline reinforcement learning
B. Kang, X. Ma, C. Du, T. Pang, and S. Yan · 2024
Closest in time.
Behavior generation with latent actions
S. Lee, Y. Wang, H. Etukuru, H. J. Kim, N. M. M. Shafiullah, and L. Pinto · 2024
Closest in time.
Discrete diffusion modeling by estimating the ratios of the data distribution
A. Lou, C. Meng, and S. Ermon · 2024
Closest in time.
Serl: A software suite for sample-efficient robotic reinforcement learning
J. Luo, Z. Hu, C. Xu, Y. L. Tan, J. Berg, A. Sharma, S. Schaal, C. Finn, A. Gupta, and S. Levine · 2024
Closest in time.
Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning
M. Nakamoto, S. Zhai, A. Singh, M. Sobol Mark, Y. Ma, C. Finn, A. Kumar, and S. Levine · 2024
Closest in time.
Simple and effective masked diffusion language models
S. S. Sahoo, M. Arriola, Y. Schiff, A. Gokaslan, E. Marroquin, J. T. Chiu, A. Rush, and V. Kuleshov · 2024
Closest in time.
Reconciling reality through simulation: A real-to-sim-to-real approach for robust manipulation
M. Torne, A. Simeonov, Z. Li, A. Chan, T. Chen, A. Gupta, and P. Agrawal · 2024
Closest in time.
Poco: Policy composition from and for heterogeneous robot learning
L. Wang, J. Zhao, Y. Du, E. H. Adelson, and R. Tedrake · 2024
Closest in time.
Robot fine-tuning made easy: Pre-training rewards and policies for autonomous real-world reinforcement learning
J. Yang, M. S. Mark, B. Vu, A. Sharma, J. Bohg, and C. Finn · 2024
Closest in time.
3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations
Y. Ze, G. Zhang, K. Zhang, C. Hu, M. Wang, and H. Xu · 2024
Closest in time.