Fetching the paper…
Reading the bibliography…
This tutorial provides a comprehensive survey of methods for fine-tuning diffusion models to optimize downstream reward functions.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Peng, X. B., A. Kumar, G. Zhang, and S. Levine (2019) · 1910
Earlier work this paper cites.
Behavior regularized offline reinforcement learning
Wu, Y., G. Tucker, and O. Nachum (2019) · 1911
Earlier work this paper cites.
Schr \ \backslash ” odinger bridge samplers
Bernton, E., J. Heng, A. Doucet, and P. E. Jacob (2019) · 1912
Earlier work this paper cites.
Reverse-time diffusion equation models
Anderson, B. D. (1982) · 1982
Earlier work this paper cites.
Comments on “representations of knowledge in complex systems” by u. grenander and mi miller
Besag, J. (1994) · 1994
Earlier work this paper cites.
Diffusions, Markov processes and martingales: Volume 2, Itô calculus
Rogers, L. C. G. and D. Williams (2000) · 2000
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., A. L. Maas, J. A. Bagnell, A. K. Dey, et al. (2008) · 2008
Earlier work this paper cites.
Relative entropy policy search
Peters, J., K. Mulling, and Y. Altun (2010) · 2010
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J., C. Meng, and S. Ermon (2020) · 2010
Earlier work this paper cites.
A generalized path integral control approach to reinforcement learning
Theodorou, E., J. Buchli, and S. Schaal (2010) · 2010
Earlier work this paper cites.
Riemann manifold langevin and hamiltonian monte carlo methods
Girolami, M. and B. Calderhead (2011) · 2011
Earlier work this paper cites.
Mcmc using hamiltonian dynamics
Neal, R. M. et al. (2011) · 2011
Earlier work this paper cites.
Statistical modeling and computation
Robert, C. (2014) · 2014
Earlier work this paper cites.
Taming the noise in reinforcement learning via soft updates
Fox, R., A. Pakman, and N. Tishby (2015) · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., S. Levine, P. Abbeel, M. Jordan, and P. Moritz (2015) · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., E. Weiss, N. Maheswaranathan, and S. Ganguli (2015) · 2015
Earlier work this paper cites.
Maximum entropy deep inverse reinforcement learning
Wulfmeier, M., P. Ondruska, and I. Posner (2015) · 2015
Earlier work this paper cites.
Guided cost learning: Deep inverse optimal control via policy optimization
Finn, C., S. Levine, and P. Abbeel (2016) · 2016
Earlier work this paper cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., H. Tang, P. Abbeel, and S. Levine (2017) · 2017
Earlier work this paper cites.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., M. Norouzi, K. Xu, and D. Schuurmans (2017) · 2017
Earlier work this paper cites.
A unified view of entropy-regularized markov decision processes
Neu, G., A. Jonsson, and V. Gómez (2017) · 2017
Earlier work this paper cites.
Equivalence between policy gradients and soft q-learning
Schulman, J., X. Chen, and P. Abbeel (2017) · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., F. Wolski, P. Dhariwal, A. Radford, and O. Klimov (2017) · 2017
Earlier work this paper cites.
Model predictive path integral control: From theory to parallel computation
Williams, G., A. Aldrich, and E. A. Theodorou (2017) · 2017
Earlier work this paper cites.
Neural ordinary differential equations
Chen, R. T., Y. Rubanova, J. Bettencourt, and D. K. Duvenaud (2018) · 2018
Earlier work this paper cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Levine, S. (2018) · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and A. G. Barto (2018) · 2018
Earlier work this paper cites.
Reinforcement learning: Theory and algorithms
Agarwal, A., N. Jiang, S. M. Kakade, and W. Sun (2019) · 2019
Earlier work this paper cites.
A theory of regularized markov decision processes
Geist, M., B. Scherrer, and O. Pietquin (2019) · 2019
Earlier work this paper cites.
Optimization of molecules via deep reinforcement learning
Zhou, Z., S. Kearnes, L. Li, R. N. Zare, and P. Riley (2019) · 2019
Earlier work this paper cites.
Controlled sequential monte carlo
Heng, J., A. N. Bishop, G. Deligiannidis, and A. Doucet (2020) · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., A. Jain, and P. Abbeel (2020) · 2020
Earlier work this paper cites.
Conservative q-learning for offline reinforcement learning
Kumar, A., A. Zhou, G. Tucker, and S. Levine (2020) · 2020
Earlier work this paper cites.
Molecular representation: going long on fingerprints
Pattanaik, L. and C. W. Coley (2020) · 2020
Earlier work this paper cites.
Mopo: Model-based offline policy optimization
Yu, T., G. Thomas, L. Yu, S. Ermon, J. Y. Zou, S. Levine, C. Finn, and T. Ma (2020) · 2020
Earlier work this paper cites.
Structured denoising diffusion models in discrete state-spaces
Austin, J., D. D. Johnson, J. Ho, D. Tarlow, and R. Van Den Berg (2021) · 2021
Earlier work this paper cites.
Machine learning for designing next-generation mrna therapeutics
Castillo-Hair, S. M. and G. Seelig (2021) · 2021
Earlier work this paper cites.
Mitigating covariate shift in imitation learning via offline data without great coverage
Chang, J. D., M. Uehara, D. Sreenivas, R. Kidambi, and W. Sun (2021) · 2021
Cited alongside, same era.
Diffusion models beat gans on image synthesis
Dhariwal, P. and A. Nichol (2021) · 2021
Cited alongside, same era.
Maximum likelihood training of score-based diffusion models
Song, Y., C. Durkan, I. Murray, and S. Ermon (2021) · 2021
Cited alongside, same era.
Pessimistic model-based offline reinforcement learning under partial coverage
Uehara, M. and W. Sun (2021) · 2021
Cited alongside, same era.
Bellman-consistent pessimism for offline reinforcement learning
Xie, T., C.-A. Cheng, N. Jiang, P. Mineiro, and A. Agarwal (2021) · 2021
Cited alongside, same era.
Idql: Implicit q-learning as an actor-critic method with diffusion policies
Hansen-Estruch, P., I. Kostrikov, M. Janner, J. G. Kuba, and S. Levine (2023) · 2023
Later among the works it cites.
Reverse diffusion monte carlo
Huang, X., H. Dong, H. Yifan, Y. Ma, and T. Zhang (2023) · 2023
Later among the works it cites.
Diffusion models for black-box optimization
Krishnamoorthy, S., S. M. Mashkaria, and A. Grover (2023) · 2023
Later among the works it cites.
Aligning text-to-image models using human feedback
Lee, K., H. Liu, M. Ryu, O. Watkins, Y. Du, C. Boutilier, P. Abbeel, M. Ghavamzadeh, and S. S. Gu (2023) · 2023
Later among the works it cites.
Latent diffusion model for dna sequence generation
Li, Z., Y. Ni, T. A. B. Huygelen, A. Das, G. Xia, G.-B. Stan, and Y. Zhao (2023) · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhang, Q. and Y. Chen (2021) · 2021
Cited alongside, same era.
The genetic and biochemical determinants of mrna degradation rates in mammals
Agarwal, V. and D. R. Kelley (2022) · 2022
Cited alongside, same era.
Building normalizing flows with stochastic interpolants
Albergo, M. S. and E. Vanden-Eijnden (2022) · 2022
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback
Bai, Y., S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, et al. (2022) · 2022
Cited alongside, same era.
A continuous time framework for discrete denoising models
Campbell, A., J. Benton, V. De Bortoli, T. Rainforth, G. Deligiannidis, and A. Doucet (2022) · 2022
Cited alongside, same era.
Diffusion posterior sampling for general noisy inverse problems
Chung, H., J. Kim, M. T. Mccann, M. L. Klasky, and J. C. Ye (2022) · 2022
Cited alongside, same era.
Improving diffusion models for inverse problems using manifold constraints
Chung, H., B. Sim, D. Ryu, and J. C. Ye (2022) · 2022
Cited alongside, same era.
Later among the works it cites.
Flow matching for generative modeling
Lipman, Y., R. T. Chen, H. Ben-Hamu, M. Nickel, and M. Le (2023) · 2023
Later among the works it cites.
Discrete diffusion language modeling by estimating the ratios of the data distribution
Lou, A., C. Meng, and S. Ermon (2023) · 2023
Later among the works it cites.
Maximum entropy gflownets with soft q-learning
Mohammadpour, S., E. Bengio, E. Frejinger, and P.-L. Bacon (2023) · 2023
Later among the works it cites.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Podell, D., Z. English, K. Lacey, A. Blattmann, T. Dockhorn, J. Müller, J. Penna, and R. Rombach (2023) · 2023
Later among the works it cites.
Aligning text-to-image diffusion models with reward backpropagation
Prabhudesai, M., A. Goyal, D. Pathak, and K. Fragkiadaki (2023) · 2023
Later among the works it cites.
Generative flow networks as entropy-regularized rl
Tiapkin, D., N. Morozov, A. Naumov, and D. Vetrov (2023) · 2023
Later among the works it cites.
Conditional flow matching: Simulation-free dynamic optimal transport
Tong, A., N. Malkin, G. Huguet, Y. Zhang, J. Rector-Brooks, K. Fatras, G. Wolf, and Y. Bengio (2023) · 2023
Later among the works it cites.
Vargas, F., W. Grathwohl, and A. Doucet (2023) · 2023
Later among the works it cites.
Diffusion model alignment using direct preference optimization
Wallace, B., M. Dang, R. Rafailov, L. Zhou, A. Lou, S. Purushwalkam, S. Ermon, C. Xiong, S. Joty, and N. Naik (2023) · 2023
Later among the works it cites.
Better aligning text-to-image models with human preference
Wu, X., K. Sun, F. Zhu, R. Zhao, and H. Li (2023) · 2023
Later among the works it cites.
Diffusion models: A comprehensive survey of methods and applications
Yang, L., Z. Zhang, Y. Song, S. Hong, R. Xu, Y. Zhao, W. Zhang, B. Cui, and M.-H. Yang (2023) · 2023
Later among the works it cites.
Reward-directed conditional diffusion: Provable distribution estimation and reward improvement
Yuan, H., K. Huang, C. Ni, M. Chen, and M. Wang (2023) · 2023
Later among the works it cites.
Towards controllable diffusion models via reward-guided exploration
Zhang, H. and T. Xu (2023) · 2023
Later among the works it cites.
Diffusion models for reinforcement learning: A survey
Zhu, Z., H. Zhao, H. He, Y. Zhong, S. Zhang, Y. Yu, and W. Zhang (2023) · 2023
Later among the works it cites.
From denoising diffusions to denoising markov models
Benton, J., Y. Shi, V. De Bortoli, G. Deligiannidis, and A. Doucet (2024) · 2024
Closest in time.
Campbell, A., J. Yim, R. Barzilay, T. Rainforth, and T. Jaakkola (2024) · 2024
Closest in time.
A survey on generative diffusion models
Cao, H., C. Tan, Z. Gao, Y. Xu, G. Chen, P.-A. Heng, and S. Z. Li (2024) · 2024
Closest in time.
An overview of diffusion models: Applications, guided generation, statistical rates and optimization
Chen, M., S. Mei, J. Fan, and M. Wang (2024) · 2024
Closest in time.
Discrete probabilistic inference as control in multi-path environments
Deleu, T., P. Nouri, N. Malkin, D. Precup, and Y. Bengio (2024) · 2024
Closest in time.
Learning universal policies via text-guided video generation
Du, Y., S. Yang, B. Dai, H. Dai, O. Nachum, J. Tenenbaum, D. Schuurmans, and P. Abbeel (2024) · 2024
Closest in time.
Recent advances in path integral control for trajectory optimization: An overview in theoretical and algorithmic perspectives
Kazim, M., J. Hong, M.-G. Kim, and K.-K. K. Kim (2024) · 2024
Closest in time.
reglm: Designing realistic regulatory dna with autoregressive language models
Lal, A., D. Garfield, T. Biancalani, and G. Eraslan (2024) · 2024
Closest in time.
Designing dna with tunable regulatory activity using discrete diffusion
Sarkar, A., Z. Tang, C. Zhao, and P. Koo (2024) · 2024
Closest in time.
Diffusion schrödinger bridge matching
Shi, Y., V. De Bortoli, A. Campbell, and A. Doucet (2024) · 2024
Closest in time.
Dirichlet flow matching with applications to dna sequence design
Stark, H., B. Jing, C. Wang, G. Corso, B. Berger, R. Barzilay, and T. Jaakkola (2024) · 2024
Closest in time.
Fine-tuning of diffusion models via stochastic control: entropy regularization and beyond
Tang, W. (2024) · 2024
Closest in time.
Score-based diffusion models via stochastic differential equations–a technical tutorial
Tang, W. and H. Zhao (2024) · 2024
Closest in time.
Fine-tuning of continuous-time diffusion models as entropy-regularized control
Uehara, M., Y. Zhao, K. Black, E. Hajiramezanali, G. Scalia, N. L. Diamant, A. M. Tseng, T. Biancalani, and S. Levine (2024) · 2024
Closest in time.
Feedback efficient online fine-tuning of diffusion models
Uehara, M., Y. Zhao, K. Black, E. Hajiramezanali, G. Scalia, N. L. Diamant, A. M. Tseng, S. Levine, and T. Biancalani (2024) · 2024
Closest in time.
Uehara, M., Y. Zhao, E. Hajiramezanali, G. Scalia, G. Eraslan, A. Lal, S. Levine, and T. Biancalani (2024) · 2024
Closest in time.
Aligning protein generative models with experimental fitness via direct preference optimization
Widatalla, T., R. Rafailov, and B. Hie (2024) · 2024
Closest in time.
Adding conditional control to diffusion models with reinforcement learning
Zhao, Y., M. Uehara, G. Scalia, T. Biancalani, S. Levine, and E. Hajiramezanali (2024) · 2024
Closest in time.