Fetching the paper…
Reading the bibliography…
Generative models have made significant impacts across various domains, largely due to their ability to scale during training by increasing data, computational resources, and model size, a phenomenon characterized by the scaling laws.
Reverse-time diffusion equation models
B. D. Anderson · 1982
Earlier work this paper cites.
Online convex optimization in the bandit setting: gradient descent without a gradient
A. D. Flaxman, A. T. Kalai, and H. B. McMahan · 2004
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma · 2013
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli · 2015
Earlier work this paper cites.
Training deep nets with sublinear memory cost
T. Chen, B. Xu, C. Zhang, and C. Guestrin · 2016
Earlier work this paper cites.
An overview of gradient descent optimization algorithms
S. Ruder · 2016
Earlier work this paper cites.
Improved techniques for training gans
T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Earlier work this paper cites.
Improved precision and recall metric for assessing generative models
T. Kynkäänniemi, T. Karras, S. Laine, J. Lehtinen, and T. Aila · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Earlier work this paper cites.
Clipscore: A reference-free evaluation metric for image captioning
J. Hessel, A. Holtzman, M. Forbes, R. L. Bras, and Y. Choi · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Earlier work this paper cites.
Classifier-free diffusion guidance
J. Ho and T. Salimans · 2022
Earlier work this paper cites.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
Earlier work this paper cites.
Elucidating the design space of diffusion-based generative models
T. Karras, M. Aittala, T. Aila, and S. Laine · 2022
Earlier work this paper cites.
Medical diffusion: denoising diffusion probabilistic models for 3d medical image generation
F. Khader, G. Mueller-Franzes, S. T. Arasteh, T. Han, C. Haarburger, M. Schulze-Hagen, P. Schad, S. Engelhardt, B. Baessler, S. Foersch, et al · 2022
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
J. Li, D. Li, C. Xiong, and S. Hoi · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Earlier work this paper cites.
The effects of reward misspecification: Mapping and mitigating misaligned models
A. Pan, K. Bhatia, and J. Steinhardt · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Earlier work this paper cites.
Photorealistic text-to-image diffusion models with deep language understanding
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans, et al · 2022
Earlier work this paper cites.
Progressive distillation for fast sampling of diffusion models
T. Salimans and J. Ho · 2022
Earlier work this paper cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, et al · 2022
Earlier work this paper cites.
Diffusers: State-of-the-art diffusion models
P. von Platen, S. Patil, A. Lozhkov, P. Cuenca, N. Lambert, K. Rasul, M. Davaadorj, D. Nair, S. Paul, W. Berman, Y. Xu, S. Liu, and T. Wolf · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Cited alongside, same era.
Simple multi-dataset detection
X. Zhou, V. Koltun, and P. Krähenbühl · 2022
Cited alongside, same era.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
Building normalizing flows with stochastic interpolants
M. S. Albergo and E. Vanden-Eijnden · 2023
Cited alongside, same era.
Training diffusion models with reinforcement learning
K. Black, M. Janner, Y. Du, I. Kostrikov, and S. Levine · 2023
Cited alongside, same era.
Reno: Enhancing one-step text-to-image models through reward-based noise optimization
L. Eyring, S. Karthik, K. Roth, A. Dosovitskiy, and Z. Akata · 2024
Later among the works it cites.
Reinforcement learning for fine-tuning text-to-image diffusion models
Y. Fan, O. Watkins, Y. Du, H. Liu, M. Ryu, C. Boutilier, P. Abbeel, M. Ghavamzadeh, K. Lee, and K. Lee · 2024
Later among the works it cites.
Stream of search (sos): Learning to search in language
K. Gandhi, D. Lee, G. Grand, M. Liu, W. Cheng, A. Sharma, and N. D. Goodman · 2024
Later among the works it cites.
New desiderata for direct preference optimization
X. Hu, T. He, and D. Wipf · 2024
Later among the works it cites.
Analyzing and improving the training dynamics of diffusion models
T. Karras, M. Aittala, J. Lehtinen, J. Hellsten, T. Aila, and S. Laine · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Fan and K. Lee · 2023
Cited alongside, same era.
Scaling laws for reward model overoptimization
L. Gao, J. Schulman, and J. Hilton · 2023
Cited alongside, same era.
Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering
Y. Hu, B. Liu, J. Kasai, Y. Wang, M. Ostendorf, R. Krishna, and N. A. Smith · 2023
Cited alongside, same era.
T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation
K. Huang, K. Sun, E. Xie, Z. Li, and X. Liu · 2023
Cited alongside, same era.
S. Karthik, K. Roth, M. Mancini, and Z. Akata · 2023
Cited alongside, same era.
Pick-a-pic: An open dataset of user preferences for text-to-image generation
Y. Kirstain, A. Polyak, U. Singer, S. Matiana, J. Penna, and O. Levy · 2023
Cited alongside, same era.
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe · 2023
Cited alongside, same era.
Later among the works it cites.
Optimizing diffusion noise can serve as universal motion priors
K. Karunratanakul, K. Preechakul, E. Aksan, T. Beeler, S. Suwajanakorn, and S. Tang · 2024
Later among the works it cites.
Confidence-aware reward optimization for fine-tuning text-to-image models
K. Kim, J. Jeong, M. An, M. Ghavamzadeh, K. Dvijotham, J. Shin, and K. Lee · 2024
Later among the works it cites.
Holistic evaluation of text-to-image models
T. Lee, M. Yasunaga, C. Meng, Y. Mai, J. S. Park, A. Gupta, Y. Zhang, D. Narayanan, H. Teufel, M. Bellagente, et al · 2024
Later among the works it cites.
Process reward model with q-value rankings
W. Li and Y. Li · 2024
Later among the works it cites.
Correcting diffusion generation through resampling
Y. Liu, Y. Zhang, T. Jaakkola, and S. Chang · 2024
Later among the works it cites.
Llmscore: Unveiling the power of large language models in text-to-image synthesis evaluation
Y. Lu, X. Yang, X. Li, X. E. Wang, and W. Y. Wang · 2024
Later among the works it cites.
Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers
N. Ma, M. Goldstein, M. S. Albergo, N. M. Boffi, E. Vanden-Eijnden, and S. Xie · 2024
Later among the works it cites.
B. Na, Y. Kim, M. Park, D. Shin, W. Kang, and I.-C. Moon · 2024
Later among the works it cites.
Ditto: Diffusion inference-time t-optimization for music generation
Z. Novack, J. McAuley, T. Berg-Kirkpatrick, and N. J. Bryan · 2024
Later among the works it cites.
Movie gen: A cast of media foundation models
A. Polyak, A. Zohar, A. Brown, A. Tjandra, A. Sinha, A. Lee, A. Vyas, B. Shi, C.-Y. Ma, C.-Y. Chuang, et al · 2024
Later among the works it cites.
Not all noises are created equally: Diffusion noise selection and optimization
Z. Qi, L. Bai, H. Xiong, et al · 2024
Later among the works it cites.
Generating images of rare concepts using pre-trained diffusion models
D. Samuel, R. Ben-Ari, S. Raviv, N. Darshan, and G. Chechik · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
C. Snell, J. Lee, K. Xu, and A. Kumar · 2024
Later among the works it cites.
Dualformer: Controllable fast and slow thinking by learning with randomized reasoning traces
D. Su, S. Sukhbaatar, M. Rabbat, Y. Tian, and Q. Zheng · 2024
Later among the works it cites.
Z. Tan, X. Yang, L. Qin, M. Yang, C. Zhang, and H. Li · 2024
Later among the works it cites.
Realfill: Reference-driven generation for authentic image completion
L. Tang, N. Ruiz, Q. Chu, Y. Li, A. Holynski, D. E. Jacobs, B. Hariharan, Y. Pritch, N. Wadhwa, K. Aberman, et al · 2024
Later among the works it cites.
Visual autoregressive modeling: Scalable image generation via next-scale prediction
K. Tian, Y. Jiang, Z. Yuan, B. Peng, and L. Wang · 2024
Later among the works it cites.
Diffusion model alignment using direct preference optimization
B. Wallace, M. Dang, R. Rafailov, L. Zhou, A. Lou, S. Purushwalkam, S. Ermon, C. Xiong, S. Joty, and N. Naik · 2024
Later among the works it cites.
Revisiting text-to-image evaluation with gecko: On metrics, prompts, and human ratings
O. Wiles, C. Zhang, I. Albuquerque, I. Kajić, S. Wang, E. Bugliarello, Y. Onoe, C. Knutsen, C. Rashtchian, J. Pont-Tuset, et al · 2024
Later among the works it cites.
Monte carlo tree search boosts reasoning via iterative preference learning
Y. Xie, A. Goyal, W. Zheng, M.-Y. Kan, T. P. Lillicrap, K. Kawaguchi, and M. Shieh · 2024
Later among the works it cites.
Imagereward: Learning and evaluating human preferences for text-to-image generation
J. Xu, X. Liu, Y. Wu, Y. Tong, Q. Li, M. Ding, J. Tang, and Y. Dong · 2024
Later among the works it cites.
Using human feedback to fine-tune diffusion models without any reward model
K. Yang, J. Tao, J. Lyu, C. Ge, J. Chen, W. Shen, X. Zhu, and X. Li · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao, and K. Narasimhan · 2024
Later among the works it cites.
Language model beats diffusion – tokenizer is key to visual generation
L. Yu, J. Lezama, N. B. Gundavarapu, L. Versari, K. Sohn, D. Minnen, Y. Cheng, V. Birodkar, A. Gupta, X. Gu, A. G. Hauptmann, B. Gong, M.-H. Yang, I. Essa, D. A. Ross, and L. Jiang · 2024
Later among the works it cites.
Golden noise for diffusion models: A learning framework
Z. Zhou, S. Shao, L. Bai, Z. Xu, B. Han, and Z. Xie · 2024
Later among the works it cites.