Fetching the paper…
Reading the bibliography…
Dynamical generative models that produce samples through an iterative process, such as Flow Matching and denoising diffusion models, have seen widespread use, but there have not been many theoretically-sound methods for improving these models with reward fine-tuning.
Dynamic programming
Richard Bellman · 1957
Earlier work this paper cites.
The Mathematical Theory of Optimal Processes
L.S. Pontryagin · 1962
Earlier work this paper cites.
Path integrals and symmetry breaking for optimal control theory
H J Kappen · 2005
Earlier work this paper cites.
Linearly-solvable markov decision problems
Emanuel Todorov · 2006
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, Anind K Dey, et al · 2008
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
Pascal Vincent · 2011
Earlier work this paper cites.
Deterministic and Stochastic Optimal Control
W.H. Fleming and R.W. Rishel · 2012
Earlier work this paper cites.
Efficient rare event simulation by optimal nonequilibrium forcing
Carsten Hartmann and Christof Schütte · 2012
Earlier work this paper cites.
Optimal control as a graphical model inference problem
Hilbert J Kappen, Vicenç Gómez, and Manfred Opper · 2012
Earlier work this paper cites.
On stochastic optimal control and reinforcement learning by approximate inference
Konrad Rawlik, Marc Toussaint, and Sethu Vijayakumar · 2013
Earlier work this paper cites.
The cross-entropy method: a unified approach to combinatorial optimization, Monte-Carlo simulation and machine learning
Reuven Y Rubinstein and Dirk P Kroese · 2013
Earlier work this paper cites.
Explicit solution of relative entropy weighted control
Joris Bierkens and Hilbert J Kappen · 2014
Earlier work this paper cites.
Policy search for path integral control
Vicenç Gómez, Hilbert J Kappen, Jan Peters, and Gerhard Neumann · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2014
Earlier work this paper cites.
Applications of the cross-entropy method to importance sampling and optimal control of diffusions
Wei Zhang, Han Wang, Carsten Hartmann, Marcus Weber, and Christof Schütte · 2014
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, and Trevor Darrell · 2015
Earlier work this paper cites.
Inceptionism: Going deeper into neural networks
Alexander Mordvintsev, Christopher Olah, and Mike Tyka · 2015
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin · 2016
Earlier work this paper cites.
Deep learning approximation for stochastic control problems
Jiequn Han and Weinan E · 2016
Earlier work this paper cites.
Stable architectures for deep neural networks
Eldad Haber and Lars Ruthotto · 2017
Earlier work this paper cites.
Neural ordinary differential equations
Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud · 2018
Earlier work this paper cites.
Optimal Control Theory: Applications to Management Science and Economics
S.P. Sethi · 2018
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Yang Song and Stefano Ermon · 2019
Earlier work this paper cites.
Theoretical guarantees for sampling and inference in generative models with latent diffusions
Belinda Tzen and Maxim Raginsky · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Scalable gradients for stochastic differential equations
Xuechen Li, Ting-Kam Leonard Wong, Ricky T. Q. Chen, and David Duvenaud · 2020
Earlier work this paper cites.
Interacting particle solutions of fokker–planck equations through gradient–log–density estimation
Dimitra Maoutsa, Sebastian Reich, and Manfred Opper · 2020
Earlier work this paper cites.
Monte carlo gradient estimation in machine learning
Shakir Mohamed, Mihaela Rosca, Michael Figurnov, and Andriy Mnih · 2020
Earlier work this paper cites.
VarGrad: A low-variance gradient estimator for variational inference
Lorenz Richter, Ayman Boustati, Nikolas Nüsken, Francisco Ruiz, and Omer Deniz Akyildiz · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Cited alongside, same era.
Fine-tuning language models from human preferences
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2020
Cited alongside, same era.
Learning neural event functions for ordinary differential equations
Ricky T. Q. Chen, Brandon Amos, and Maximilian Nickel · 2021
Cited alongside, same era.
Diffusion schrödinger bridge with applications to score-based generative modeling
Valentin De Bortoli, James Thornton, Jeremy Heng, and Arnaud Doucet · 2021
Cited alongside, same era.
Clipscore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi · 2021
Cited alongside, same era.
Pick-a-pic: An open dataset of user preferences for text-to-image generation
Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana, Joe Penna, and Omer Levy · 2023
Later among the works it cites.
Flow matching for generative modeling
Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matthew Le · 2023
Later among the works it cites.
Flow straight and fast: Learning to generate and transfer data with rectified flow
Xingchao Liu, Chengyue Gong, and qiang liu · 2023
Later among the works it cites.
Trajectory balance: Improved credit assignment in gflownets
Nikolay Malkin, Moksh Jain, Emmanuel Bengio, Chen Sun, and Yoshua Bengio · 2023
Later among the works it cites.
Training-free linear image inversion via flows
Ashwini Pokle, Matthew J Muckley, Ricky T. Q. Chen, and Brian Karrer · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Openclip, July 2021
Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt · 2021
Cited alongside, same era.
On density estimation with diffusion models
Diederik P Kingma, Tim Salimans, Ben Poole, and Jonathan Ho · 2021
Cited alongside, same era.
Differentiable gaussianization layers for inverse problems regularized by deep generative models
Dongzhuo Li · 2021
Cited alongside, same era.
Solving high-dimensional Hamilton–Jacobi–Bellman pdes using neural networks: perspectives from the theory of controlled diffusions and measures on path space
Nikolas Nüsken and Lorenz Richter · 2021
Cited alongside, same era.
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan · 2022
Cited alongside, same era.
Diffusion posterior sampling for general noisy inverse problems
Hyungjin Chung, Jeongsol Kim, Michael T Mccann, Marc L Klasky, and Jong Chul Ye · 2022
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2023
Later among the works it cites.
Pseudoinverse-guided diffusion models for inverse problems
Jiaming Song, Arash Vahdat, Morteza Mardani, and Jan Kautz · 2023
Later among the works it cites.
Denoising diffusion samplers
Francisco Vargas, Will Sussman Grathwohl, and Arnaud Doucet · 2023
Later among the works it cites.
Audiobox: Unified audio generation with natural language prompts
Apoorv Vyas, Bowen Shi, Matthew Le, Andros Tjandra, Yi-Chiao Wu, Baishan Guo, Jiemin Zhang, Xinyue Zhang, Robert Adkins, William Ngan, et al · 2023
Later among the works it cites.
Imagereward: Learning and evaluating human preferences for text-to-image generation
Jiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong, Qinkai Li, Ming Ding, Jie Tang, and Yuxiao Dong · 2023
Later among the works it cites.
Guided flows for generative modeling and decision making
Qinqing Zheng, Matt Le, Neta Shaul, Yaron Lipman, Aditya Grover, and Ricky T. Q. Chen · 2023
Later among the works it cites.
Consistency-diversity-realism pareto fronts of conditional image generative models
Pietro Astolfi, Marlene Careil, Melissa Hall, Oscar Mañas, Matthew Muckley, Jakob Verbeek, Adriana Romero Soriano, and Michal Drozdzal · 2024
Closest in time.
D-flow: Differentiating through flows for controlled generation
Heli Ben-Hamu, Omri Puny, Itai Gat, Brian Karrer, Uriel Singer, and Yaron Lipman · 2024
Closest in time.
Training diffusion models with reinforcement learning
Kevin Black, Michael Janner, Yilun Du, Ilya Kostrikov, and Sergey Levine · 2024
Closest in time.
Posterior sampling with denoising oracles via tilted transport
Joan Bruna and Jiequn Han · 2024
Closest in time.
Directly fine-tuning diffusion models on differentiable rewards
Kevin Clark, Paul Vicol, Kevin Swersky, and David J. Fleet · 2024
Closest in time.
Alexander Denker, Francisco Vargas, Shreyas Padhy, Kieran Didi, Simon Mathis, Vincent Dutordoir, Riccardo Barbano, Emile Mathieu, Urszula Julia Komorowska, and Pietro Lio · 2024
Closest in time.
A taxonomy of loss functions for stochastic optimal control
Carles Domingo-Enrich · 2024
Closest in time.
Doob’s lagrangian: A sample-efficient variational approach to transition path sampling
Yuanqi Du, Michael Plainer, Rob Brekelmans, Chenru Duan, Frank Noé, Carla P. Gomes, Alan Apsuru-Guzik, and Kirill Neklyudov · 2024
Closest in time.
Scaling rectified flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al · 2024
Closest in time.
Reinforcement learning for fine-tuning text-to-image diffusion models
Ying Fan, Olivia Watkins, Yuqing Du, Hao Liu, Moonkyung Ryu, Craig Boutilier, Pieter Abbeel, Mohammad Ghavamzadeh, Kangwook Lee, and Kimin Lee · 2024
Closest in time.
Symbolic music generation with non-differentiable rule guided diffusion
Yujia Huang, Adishree Ghatare, Yuanzhe Liu, Ziniu Hu, Qinsheng Zhang, Chandramouli S Sastry, Siddharth Gururani, Sageev Oore, and Yisong Yue · 2024
Closest in time.
Voicebox: Text-guided multilingual universal speech generation at scale
Matthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer, Leda Sari, Rashel Moritz, Mary Williamson, Vimal Manohar, Yossi Adi, Jay Mahadeokar, et al · 2024
Closest in time.
Generalized schrödinger bridge matching
Guan-Horng Liu, Yaron Lipman, Maximilian Nickel, Brian Karrer, Evangelos Theodorou, and Ricky T. Q. Chen · 2024
Closest in time.
Neural optimal transport with lagrangian costs
Aram-Alexandre Pooladian, Carles Domingo-Enrich, Ricky T. Q. Chen, and Brandon Amos · 2024
Closest in time.
Improved sampling via learned diffusions
Lorenz Richter and Julius Berner · 2024
Closest in time.
Rb-modulation: Training-free personalization of diffusion models using stochastic optimal control
Litu Rout, Yujia Chen, Nataniel Ruiz, Abhishek Kumar, Constantine Caramanis, Sanjay Shakkottai, and Wen-Sheng Chu · 2024
Closest in time.
Fine-tuning of diffusion models via stochastic control: entropy regularization and beyond
Wenpin Tang · 2024
Closest in time.
Improving gflownets for text-to-image diffusion alignment
Dinghuai Zhang, Yizhe Zhang, Jiatao Gu, Ruixiang Zhang, Josh Susskind, Navdeep Jaitly, and Shuangfei Zhai · 2024
Closest in time.