Fetching the paper…
Reading the bibliography…
Understanding visual scenes is fundamental to human intelligence.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Accurately computing the log-sum-exp and softmax functions
Pierre Blanchard, Desmond J Higham, and Nicholas J Higham · 2021
Earlier work this paper cites.
Openclip, July 2021
Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Classifier-Free Diffusion Guidance, July 2022
Jonathan Ho and Tim Salimans · 2022
Earlier work this paper cites.
Elucidating the design space of diffusion-based generative models
Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine · 2022
Earlier work this paper cites.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Earlier work this paper cites.
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al · 2022
Earlier work this paper cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al · 2022
Earlier work this paper cites.
Winoground: Probing vision and language models for visio-linguistic compositionality
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross · 2022
Earlier work this paper cites.
Building normalizing flows with stochastic interpolants
Michael S Albergo and Eric Vanden-Eijnden · 2023
Earlier work this paper cites.
Stable video diffusion: Scaling latent video diffusion models to large datasets
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, et al · 2023
Earlier work this paper cites.
Pixart-alpha: Fast training of diffusion transformer for photorealistic text-to-image synthesis
Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, et al · 2023
Earlier work this paper cites.
Text-to-image diffusion models are zero shot classifiers
Kevin Clark and Priyank Jaini · 2023
Earlier work this paper cites.
Geneval: An object-focused framework for evaluating text-to-image alignment
Dhruba Ghosh, Hannaneh Hajishirzi, and Ludwig Schmidt · 2023
Earlier work this paper cites.
Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality
Cheng-Yu Hsieh, Jieyu Zhang, Zixian Ma, Aniruddha Kembhavi, and Ranjay Krishna · 2023
Cited alongside, same era.
T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation
Kaiyi Huang, Kaiyue Sun, Enze Xie, Zhenguo Li, and Xihui Liu · 2023
Cited alongside, same era.
Midjourney
Midjourney Inc · 2023
Cited alongside, same era.
Intriguing properties of generative classifiers
Priyank Jaini, Kevin Clark, and Robert Geirhos · 2023
Cited alongside, same era.
The power of sound (tpos): Audio reactive video generation with stable diffusion
Yujin Jeong, Wonjeong Ryoo, Seunghyun Lee, Dabin Seo, Wonmin Byeon, Sangpil Kim, and Jinkyu Kim · 2023
Cited alongside, same era.
What’s" up" with vision-language models? investigating their struggle with spatial reasoning
Fast timing-conditioned latent audio diffusion
Zach Evans, CJ Carr, Josiah Taylor, Scott H Hawley, and Jordi Pons · 2024
Later among the works it cites.
Discffusion: Discriminative diffusion models as few-shot vision and language learners
Xuehai He, Weixi Feng, Tsu-Jui Fu, Varun Jampani, Arjun Akula, Pradyumna Narayana, Sugato Basu, William Yang Wang, and Xin Eric Wang · 2024
Later among the works it cites.
Adaptive non-uniform timestep sampling for diffusion model training, 2024
Myunsoo Kim, Donghyeon Ki, Seong-Woong Shim, and Byung-Jun Lee · 2024
Later among the works it cites.
Understanding diffusion-based representation learning via low-dimensional modeling
Xiao Li, Zekai Zhang, Xiang Li, Siyi Chen, Zhihui Zhu, Peng Wang, and Qing Qu · 2024
Later among the works it cites.
Emergence of hidden capabilities: Exploring learning dynamics in concept space
Core Francisco Park, Maya Okawa, Andrew Lee, Ekdeep S Lubana, and Hidenori Tanaka · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Amita Kamath, Jack Hessel, and Kai-Wei Chang · 2023
Cited alongside, same era.
Are diffusion models vision-and-language reasoners?
Benno Krojer, Elinor Poole-Dayan, Vikram Voleti, Christopher Pal, and Siva Reddy · 2023
Cited alongside, same era.
Your diffusion model is secretly a zero-shot classifier
Alexander C Li, Mihir Prabhudesai, Shivam Duggal, Ellis Brown, and Deepak Pathak · 2023
Cited alongside, same era.
Flow matching for generative modeling
Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le · 2023
Cited alongside, same era.
Teaching clip to count to ten
Roni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada, Inbar Mosseri, Michal Irani, and Tali Dekel · 2023
Cited alongside, same era.
Cola: A benchmark for compositional text-to-image retrieval
Arijit Ray, Filip Radenovic, Abhimanyu Dubey, Bryan Plummer, Ranjay Krishna, and Kate Saenko · 2023
Cited alongside, same era.
The surprising effectiveness of diffusion models for optical flow and monocular depth estimation
Saurabh Saxena, Charles Herrmann, Junhwa Hur, Abhishek Kar, Mohammad Norouzi, Deqing Sun, and David J Fleet · 2023
Cited alongside, same era.
Synthesize diagnose and optimize: Towards fine-grained vision-language understanding
Wujian Peng, Sicheng Xie, Zuyao You, Shiyi Lan, and Zuxuan Wu · 2024
Later among the works it cites.
A simple and efficient baseline for zero-shot generative classification
Zipeng Qi, Buhua Liu, Shiyan Zhang, Bao Li, Zhiqiang Xu, Haoyi Xiong, and Zeke Xie · 2024
Later among the works it cites.
Adversarial diffusion distillation
Axel Sauer, Dominik Lorenz, Andreas Blattmann, and Robin Rombach · 2024
Later among the works it cites.
Eyes wide shut? exploring the visual shortcomings of multimodal llms
Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, and Saining Xie · 2024
Later among the works it cites.
Omnigen: Unified image generation
Shitao Xiao, Yueze Wang, Junjie Zhou, Huaying Yuan, Xingrun Xing, Ruiran Yan, Chaofan Li, Shuting Wang, Tiejun Huang, and Zheng Liu · 2024
Later among the works it cites.
Chenglin Yang, Celong Liu, Xueqing Deng, Dongwon Kim, Xing Mei, Xiaohui Shen, and Liang-Chieh Chen · 2024
Later among the works it cites.
Representation alignment for generation: Training diffusion transformers is easier than you think
Sihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong, Jonathan Huang, Jinwoo Shin, and Saining Xie · 2024
Later among the works it cites.
Few-shot learner parameterization by diffusion time-steps
Zhongqi Yue, Pan Zhou, Richang Hong, Hanwang Zhang, and Qianru Sun · 2024
Later among the works it cites.
Beyond overcorrection: Evaluating diversity in t2i models with divbench
Felix Friedrich, Thiemo Ganesha Welsch, Manuel Brack, Patrick Schramowski, and Kristian Kersting · 2025
Closest in time.
Read, watch and scream! sound generation from text and video
Yujin Jeong, Yunji Kim, Sanghyuk Chun, and Jiyoung Lee · 2025
Closest in time.
Clip behaves like a bag-of-words model cross-modally but not uni-modally
Darina Koishigarina, Arnas Uselis, and Seong Joon Oh · 2025
Closest in time.
Stai-tuned experiment scheduler: Structured experiment management with google sheets, weights & biases, and hpc integration
Alexander Rubinstein and Arnas Uselis · 2025
Closest in time.
Intermediate layer classifiers for ood generalization, 2025
Arnas Uselis and Seong Joon Oh · 2025
Closest in time.
Does data scaling lead to visual compositional generalization?
Arnas Uselis, Andrea Dittadi, and Seong Joon Oh · 2025
Closest in time.
A closer look at time steps is worthy of triple speed-up for diffusion model training, 2025
Kai Wang, Mingjia Shi, Yukun Zhou, Zekai Li, Zhihang Yuan, Yuzhang Shang, Xiaojiang Peng, Hanwang Zhang, and Yang You · 2025
Closest in time.