Fetching the paper…
Reading the bibliography…
Generative AI has made remarkable strides to revolutionize fields such as image and video generation.
Rank analysis of incomplete block designs: I. the method of paired comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
The behavior of maximum likelihood estimates under nonstandard conditions
P. J. Huber et al · 1967
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli · 2004
Earlier work this paper cites.
Improved techniques for training gans
T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, X. Chen, and X. Chen · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
T. Unterthiner, S. van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang · 2018
Earlier work this paper cites.
CLIPScore: a reference-free evaluation metric for image captioning
J. Hessel, A. Holtzman, M. Forbes, R. L. Bras, and Y. Choi · 2021
Earlier work this paper cites.
Sdedit: Guided image synthesis and editing with stochastic differential equations
C. Meng, Y. He, Y. Song, J. Song, J. Wu, J.-Y. Zhu, and S. Ermon · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
Prompt-to-prompt image editing with cross attention control
A. Hertz, R. Mokady, J. M. Tenenbaum, K. Aberman, Y. Pritch, and D. Cohen-Or · 2022
Earlier work this paper cites.
Imagen video: High definition video generation with diffusion models
J. Ho, W. Chan, C. Saharia, J. Whang, R. Gao, A. Gritsenko, D. P. Kingma, B. Poole, M. Norouzi, D. J. Fleet, et al · 2022
Earlier work this paper cites.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
A. Q. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. Mcgrew, I. Sutskever, and M. Chen · 2022
Earlier work this paper cites.
Hierarchical text-conditional image generation with clip latents
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Earlier work this paper cites.
Photorealistic text-to-image diffusion models with deep language understanding
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans, et al · 2022
Earlier work this paper cites.
Qwen-vl: A frontier large vision-language model with versatile abilities
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou · 2023
Earlier work this paper cites.
Introducing our multimodal models, 2023
R. Bavishi, E. Elsen, C. Hawthorne, M. Nye, A. Odena, A. Somani, and S. Taşırlar · 2023
Earlier work this paper cites.
Stable video diffusion: Scaling latent video diffusion models to large datasets
A. Blattmann, T. Dockhorn, S. Kulal, D. Mendelevitch, M. Kilian, and D. Lorenz · 2023
Earlier work this paper cites.
Instructpix2pix: Learning to follow image editing instructions
T. Brooks, A. Holynski, and A. A. Efros · 2023
Earlier work this paper cites.
Pixart- α \alpha : Fast training of diffusion transformer for photorealistic text-to-image synthesis
J. Chen, J. Yu, C. Ge, L. Yao, E. Xie, Y. Wu, Z. Wang, J. T. Kwok, P. Luo, H. Lu, and Z. Li · 2023
Earlier work this paper cites.
Lipsim: A provably robust perceptual similarity metric
S. Ghazanfari, A. Araujo, P. Krishnamurthy, F. Khorrami, and S. Garg · 2023
Earlier work this paper cites.
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Y. Guo, C. Yang, A. Rao, Z. Liang, Y. Wang, Y. Qiao, M. Agrawala, D. Lin, and B. Dai · 2023
Cited alongside, same era.
Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering
Y. Hu, B. Liu, J. Kasai, Y. Wang, M. Ostendorf, R. Krishna, and N. A. Smith · 2023
Cited alongside, same era.
T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation
K. Huang, K. Sun, E. Xie, Z. Li, and X. Liu · 2023
Cited alongside, same era.
Llama guard: Llm-based input-output safeguard for human-ai conversations
H. Inan, K. Upasani, J. Chi, R. Rungta, K. Iyer, Y. Mao, M. Tontchev, Q. Hu, B. Fuller, D. Testuggine, and M. Khabsa · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica · 2023
Minigpt-4: Enhancing vision-language understanding with advanced large language models
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny · 2023
Later among the works it cites.
Auraflow, 2024
fal · 2024
Closest in time.
Dreamsim: Learning new dimensions of human visual similarity using synthetic data
S. Fu, N. Tamir, S. Sundaram, L. Chai, R. Zhang, T. Dekel, and P. Isola · 2024
Closest in time.
Minicpm: Unveiling the potential of small language models with scalable training strategies
S. Hu, Y. Tu, X. Han, C. He, G. Cui, X. Long, Z. Zheng, Y. Fang, Y. Huang, W. Zhao, X. Zhang, Z. L. Thai, K. Zhang, C. Wang, Y. Yao, C. Zhao, J. Zhou, J. Cai, Z. Zhai, N. Ding, C. Jia, G. Zeng, D. Li, Z. Liu, and M. Sun · 2024
Closest in time.
VBench: Comprehensive benchmark suite for video generative models
Z. Huang, Y. He, J. Yu, F. Zhang, C. Si, Y. Jiang, Y. Zhang, T. Wu, Q. Jin, N. Chanpaisit, Y. Wang, X. Chen, L. Wang, D. Lin, Y. Qiao, and Z. Liu · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Obelics: An open web-scale filtered dataset of interleaved image-text documents, 2023
H. Laurençon, L. Saulnier, L. Tronchon, S. Bekman, A. Singh, A. Lozhkov, T. Wang, S. Karamcheti, A. M. Rush, D. Kiela, M. Cord, and V. Sanh · 2023
Cited alongside, same era.
Video-llava: Learning united visual representation by alignment before projection
B. Lin, B. Zhu, Y. Ye, M. Ning, P. Jin, and L. Yuan · 2023
Cited alongside, same era.
Latent consistency models: Synthesizing high-resolution images with few-step inference
S. Luo, Y. Tan, L. Huang, J. Li, and H. Zhao · 2023
Cited alongside, same era.
Gpt-4 technical report, 2023
OpenAI · 2023
Cited alongside, same era.
Openjourney is an open source stable diffusion fine tuned model on midjourney images, 2023
openjourney.ai · 2023
Cited alongside, same era.
Toward verifiable and reproducible human evaluation for text-to-image generation
M. Otani, R. Togashi, Y. Sawai, R. Ishigami, Y. Nakashima, E. Rahtu, J. Heikkilä, and S. Satoh · 2023
Cited alongside, same era.
Zero-shot image-to-image translation
G. Parmar, K. Kumar Singh, R. Zhang, Y. Li, J. Lu, and J.-Y. Zhu · 2023
Cited alongside, same era.
Tokenizer arena
Hugging Face Spaces · 2024
Closest in time.
Mantis: Interleaved multi-image instruction tuning
D. Jiang, X. He, H. Zeng, C. Wei, M. Ku, Q. Liu, and W. Chen · 2024
Closest in time.
Kolors, 2024
Kwai-Kolors · 2024
Closest in time.
Flux, 2024
B. F. Labs · 2024
Closest in time.
What matters when building vision-language models?, 2024
H. Laurençon, L. Tronchon, M. Cord, and V. Sanh · 2024
Closest in time.
Holistic evaluation of text-to-image models
T. Lee, M. Yasunaga, C. Meng, Y. Mai, J. S. Park, A. Gupta, Y. Zhang, D. Narayanan, H. Teufel, M. Bellagente, et al · 2024
Closest in time.
Sdxl-lightning: Progressive adversarial diffusion distillation
S. Lin, A. Wang, and X. Yang · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee · 2024
Closest in time.
Text to speech arena
mrfakename, V. Srivastav, C. Fourrier, L. Pouget, Y. Lacombe, main, and S. Gandhi · 2024
Closest in time.
Open-Sora: Democratizing Efficient Video Production for All
N. U. of Singapore · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
M. Reid, N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J.-b. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser, et al · 2024
Closest in time.
Stable diffusion 3 release, 2024
Stability AI · 2024
Closest in time.
Inversion-free image editing with natural language
S. Xu, Y. Huang, J. Pan, Z. Ma, and J. Chai · 2024
Closest in time.
Cogvideox: Text-to-video diffusion models with an expert transformer
Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y. Yang, W. Hong, X. Zhang, G. Feng, et al · 2024
Closest in time.
Lefusion: Synthesizing myocardial pathology on cardiac mri via lesion-focus diffusion models, 2024
H. Zhang, J. Yang, S. Wan, and P. Fua · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al · 2024
Closest in time.