Fetching the paper…
Reading the bibliography…
Recent advancements in image generative foundation models have prioritized quality improvements but often at the cost of increased computational complexity and inference latency.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke · 2019
Earlier work this paper cites.
Glu variants improve transformer
N. Shazeer · 2020
Earlier work this paper cites.
Flow matching for generative modeling
Y. Lipman, R. T. Chen, H. Ben-Hamu, M. Nickel, and M. Le · 2022
Earlier work this paper cites.
A self-supervised descriptor for image copy detection
E. Pizzi, S. D. Roy, S. N. Ravindra, P. Goyal, and M. Douze · 2022
Earlier work this paper cites.
Geneval: An object-focused framework for evaluating text-to-image alignment
D. Ghosh, H. Hajishirzi, and L. Schmidt · 2023
Earlier work this paper cites.
Diffusion art or digital forgery? investigating data replication in diffusion models
G. Somepalli, V. Singla, M. Goldblum, J. Geiping, and T. Goldstein · 2023
Earlier work this paper cites.
X. Wu, Y. Hao, K. Sun, Y. Chen, F. Zhu, R. Zhao, and H. Li · 2023
Cited alongside, same era.
Magicbrush: A manually annotated dataset for instruction-guided image editing
K. Zhang, L. Mo, W. Chen, H. Sun, and Y. Su · 2023
Cited alongside, same era.
URL https://blackforestlabs.ai/
Black forest labs, 2024 · 2024
Cited alongside, same era.
Topiq: A top-down approach from semantics to distortions for image quality assessment
C. Chen, J. Mo, J. Hou, H. Wu, L. Liao, W. Sun, Q. Yan, and W. Lin · 2024
Cited alongside, same era.
M. Douze, A. Guzhva, C. Deng, J. Johnson, G. Szilvasy, P.-E. Mazaré, M. Lomeli, L. Hosseini, and H. Jégou · 2024
Cited alongside, same era.
Smartedit: Exploring complex instruction-based image editing with multimodal large language models
Y. Huang, L. Xie, X. Wang, Z. Yuan, X. Cun, Y. Ge, J. Zhou, C. Dong, R. Huang, R. Zhang, et al · 2024
Later among the works it cites.
Minicpm-v 2.6: A high-performance vision language model
OpenBMB · 2024
Later among the works it cites.
Emu edit: Precise image editing via recognition and generation tasks
S. Sheynin, A. Polyak, U. Singer, Y. Kirstain, A. Zohar, O. Ashual, D. Parikh, and Y. Taigman · 2024
Later among the works it cites.
Omnigen: Unified image generation
S. Xiao, Y. Wang, J. Zhou, H. Yuan, X. Xing, R. Yan, C. Li, S. Wang, T. Huang, and Z. Liu · 2024
Later among the works it cites.
Improved distribution matching distillation for fast image synthesis
T. Yin, M. Gharbi, T. Park, R. Zhang, E. Shechtman, F. Durand, and B. Freeman · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scaling rectified flow transformers for high-resolution image synthesis
P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, H. Saini, Y. Levi, D. Lorenz, A. Sauer, F. Boesel, et al · 2024
Cited alongside, same era.
Ella: Equip diffusion models with llm for enhanced semantic alignment
X. Hu, R. Wang, Y. Fang, B. Fu, P. Cheng, and G. Yu · 2024
Cited alongside, same era.
Aesthetic predictor
LAION-AI
Cited in the paper.
Clip-based nsfw detector
LAION-AI
Cited in the paper.
Laion-5b-watermarkdetection
LAION-AI
Cited in the paper.
Long-clip: Unlocking the long-text capability of clip
B. Zhang, P. Zhang, X. Dong, Y. Zang, and J. Wang · 2024
Later among the works it cites.
Ultraedit: Instruction-based fine-grained image editing at scale
H. Zhao, X. S. Ma, L. Chen, S. Si, R. Wu, K. An, P. Yu, M. Zhang, Q. Li, and B. Chang · 2024
Later among the works it cites.