Fetching the paper…
Reading the bibliography…
Diffusion policies have achieved superior performance in imitation learning and offline reinforcement learning (RL) due to their rich expressiveness.
Estimation of non-normalized statistical models by score matching
Hyvärinen, A. and Dayan, P · 2005
Earlier work this paper cites.
Mirror descent policy optimization, 2021
Tomar, M., Shani, L., Efroni, Y., and Ghavamzadeh, M · 2005
Earlier work this paper cites.
Relative entropy policy search
Peters, J., Mulling, K., and Altun, Y · 2010
Earlier work this paper cites.
Tweedie’s formula and selection bias
Efron, B · 2011
Earlier work this paper cites.
Mcmc using hamiltonian dynamics
Neal, R. M. et al · 2011
Earlier work this paper cites.
A Connection Between Score Matching and Denoising Autoencoders
Vincent, P · 2011
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
Trust region policy optimization
Schulman, J · 2015
Earlier work this paper cites.
Deep Unsupervised Learning using Nonequilibrium Thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Reinforcement Learning with Deep Energy-Based Policies, July 2017
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D · 2017
Earlier work this paper cites.
A unified view of entropy-regularized markov decision processes
Neu, G., Jonsson, A., and Gómez, V · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., and Meger, D · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Song, Y. and Ermon, S · 2019
Cited alongside, same era.
Denoising Diffusion Probabilistic Models
Ho, J., Jain, A., and Abbeel, P · 2020
Cited alongside, same era.
On the global convergence rates of softmax policy gradient methods
Mei, J., Xiao, C., Szepesvari, C., and Schuurmans, D · 2020
Cited alongside, same era.
Sliced score matching: A scalable approach to density and score estimation
Song, Y., Garg, S., Shi, J., and Ermon, S · 2020
Cited alongside, same era.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Cited alongside, same era.
How to train your energy-based models
Song, Y. and Kingma, D. P · 2021
Diffusion policy: Visuomotor policy learning via action diffusion
Chi, C., Xu, Z., Feng, S., Cousineau, E., Du, Y., Burchfiel, B., Tedrake, R., and Song, S · 2023
Later among the works it cites.
Idql: Implicit q-learning as an actor-critic method with diffusion policies
Hansen-Estruch, P., Kostrikov, I., Janner, M., Kuba, J. G., and Levine, S · 2023
Later among the works it cites.
Policy mirror descent for reinforcement learning: Linear convergence, new sampling complexity, and generalized problem classes
Lan, G · 2023
Later among the works it cites.
Learning a diffusion model policy from rewards via q-score matching
Psenka, M., Escontrela, A., Abbeel, P., and Ma, Y · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Score-Based Generative Modeling through Stochastic Differential Equations, February 2021
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2021
Cited alongside, same era.
Offline reinforcement learning via high-fidelity generative behavior modeling
Chen, H., Lu, C., Ying, C., Su, H., and Zhu, J · 2022
Cited alongside, same era.
Planning with diffusion for flexible behavior synthesis
Janner, M., Du, Y., Tenenbaum, J. B., and Levine, S · 2022
Cited alongside, same era.
Elucidating the Design Space of Diffusion-Based Generative Models, October 2022
Karras, T., Aittala, M., Aila, T., and Laine, S · 2022
Cited alongside, same era.
Flow annealed importance sampling bootstrap
Midgley, L. I., Stimper, V., Simm, G. N., Schölkopf, B., and Hernández-Lobato, J. M · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al · 2022
Cited alongside, same era.
Rigter, M., Yamada, J., and Posner, I · 2023
Later among the works it cites.
Iterated denoising energy matching for sampling from boltzmann densities
Akhound-Sadegh, T., Rector-Brooks, J., Bose, A. J., Mittal, S., Lemos, P., Liu, C.-H., Sendera, M., Ravanbakhsh, S., Gidel, G., Bengio, Y., et al · 2024
Later among the works it cites.
Learning universal policies via text-guided video generation
Du, Y., Yang, S., Dai, B., Dai, H., Nachum, O., Tenenbaum, J., Schuurmans, D., and Abbeel, P · 2024
Later among the works it cites.
Diffuseloco: Real-time legged locomotion control with diffusion from offline datasets
Huang, X., Chi, Y., Wang, R., Li, Z., Peng, X. B., Shao, S., Nikolic, B., and Sreenath, K · 2024
Later among the works it cites.
Sampling from energy-based policies using diffusion
Jain, V., Akhound-Sadegh, T., and Ravanbakhsh, S · 2024
Later among the works it cites.
3d diffuser actor: Policy diffusion with 3d scene representations
Ke, T.-W., Gkanatsios, N., and Fragkiadaki, K · 2024
Later among the works it cites.
Diffusion policy policy optimization
Ren, A. Z., Lidard, J., Ankile, L. L., Simeonov, A., Agrawal, P., Majumdar, A., Burchfiel, B., Dai, H., and Simchowitz, M · 2024
Later among the works it cites.
Movement primitive diffusion: Learning gentle robotic manipulation of deformable objects
Scheikl, P. M., Schreiber, N., Haas, C., Freymuth, N., Neumann, G., Lioutikov, R., and Mathis-Ullrich, F · 2024
Later among the works it cites.
Diffusion spectral representation for reinforcement learning
Shribak, D., Gao, C.-X., Li, Y., Xiao, C., and Dai, B · 2024
Later among the works it cites.
Diffusion Actor-Critic with Entropy Regulator, December 2024
Wang, Y., Wang, L., Jiang, Y., Zou, W., Liu, T., Song, X., Wang, W., Xiao, L., Wu, J., Duan, J., and Li, S. E · 2024
Later among the works it cites.
Dime: Diffusion-based maximum entropy reinforcement learning
Celik, O., Li, Z., Blessing, D., Li, G., Palanicek, D., Peters, J., Chalvatzaki, G., and Neumann, G · 2025
Closest in time.