2023

Reward-Directed Conditional Diffusion: Provable Distribution Estimation and Reward Improvement

Yuan, Hui, Huang, Kaixuan, Ni, Chengzhuo et al.

Understand

We explore the methodology and theory of reward-directed generation via conditional diffusion models.

  • Directed generation aims to generate samples with desired properties as measured by a reward function, which has broad applications in generative AI, reinforcement learning, and computational biology.
  • We consider the common learning scenario where the data set consists of unlabeled data along with a smaller set of data with noisy reward labels.
  • Our approach leverages a learned reward function on the smaller data set as a pseudolabeler.

Reading the bibliography…