Fetching the paper…
Reading the bibliography…
The burgeoning landscape of text-to-image models, exemplified by innovations such as Midjourney and DALLE 3, has revolutionized content creation across diverse sectors.
Quantifying the carbon emissions of machine learning
Lacoste, A., Luccioni, A., Schmidt, V., and Dandres, T. (2019) · 1910
Earlier work this paper cites.
I, robot vol. 1
Asimov, I · 2004
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L. (2014) · 2014
Earlier work this paper cites.
arXiv preprint arXiv:1412.6980
Kingma, D. P., and Ba, J. (2014). Adam: A method for stochastic optimization · 2014
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015) · 2015
Earlier work this paper cites.
Advances in neural information processing systems 29
Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., and Chen, X. (2016). Improved techniques for training gans · 2016
Earlier work this paper cites.
Advances in neural information processing systems 30
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S. (2017). Gans trained by a two time-scale update rule converge to a local nash equilibrium · 2017
Earlier work this paper cites.
Md. L. Rev. 78 , 892
Franks, M. A., and Waldman, A. E. (2018). Sex, lies, and videotape: Deep fakes and free speech delusions · 2018
Earlier work this paper cites.
pytorch-fid: FID Score for PyTorch
Seitzer, M. (2020) · 2020
Earlier work this paper cites.
High-fidelity performance metrics for generative models in pytorch
Obukhov, A., Seitzer, M., Wu, P.-W., Zhydenko, S., Kyl, J., and Lin, E. Y.-J. (2020) · 2020
Earlier work this paper cites.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I. (2021) · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. (2021) · 2021
Earlier work this paper cites.
Fairface: Face attribute dataset for balanced race, gender, and age for bias measurement and mitigation
Karkkainen, K., and Joo, J. (2021) · 2021
Earlier work this paper cites.
arXiv preprint arXiv:2106.09685
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. (2021). Lora: Low-rank adaptation of large language models · 2021
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. (2022) · 2022
Earlier work this paper cites.
Make-a-scene: Scene-based text-to-image generation with human priors
Gafni, O., Polyak, A., Ashual, O., Sheynin, S., Parikh, D., and Taigman, Y. (2022) · 2022
Earlier work this paper cites.
Scaling autoregressive models for content-rich text-to-image generation
Yu, J., Xu, Y., Koh, J. Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B. K., Hutchinson, B., Han, W., Parekh, Z., Li, X., Zhang, H., Baldridge, J., and Wu, Y. (2022) · 2022
Earlier work this paper cites.
arXiv preprint arXiv:2211.12561
Yasunaga, M., Aghajanyan, A., Shi, W., James, R., Leskovec, J., Liang, P., Lewis, M., Zettlemoyer, L., and Yih, W.-t. (2022). Retrieval-augmented multimodal language modeling · 2022
Cited alongside, same era.
https://www.midjourney.com/
Midjourney (2022) · 2022
Cited alongside, same era.
arXiv preprint arXiv:2204.06125
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. (2022). Hierarchical text-conditional image generation with clip latents. arxiv 2022 · 2022
Cited alongside, same era.
Unstable diffusion: Ethical challenges and some ways forward
Gupta, A. (2022) · 2022
Cited alongside, same era.
Image segmentation using text and image prompts
Lüddecke, T., and Ecker, A. S. (2022) · 2022
Cited alongside, same era.
arXiv preprint arXiv:2302.10893
Friedrich, F., Schramowski, P., Brack, M., Struppek, L., Hintersdorf, D., Luccioni, S., and Kersting, K. (2023). Fair diffusion: Instructing text-to-image generation models on fairness · 2023
Later among the works it cites.
Holistic evaluation of text-to-image models
Lee, T., Yasunaga, M., Meng, C., Mai, Y., Park, J. S., Gupta, A., Zhang, Y., Narayanan, D., Teufel, H. B., Bellagente, M., Kang, M., Park, T., Leskovec, J., Zhu, J.-Y., Fei-Fei, L., Wu, J., Ermon, S., and Liang, P. (2023) · 2023
Later among the works it cites.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R. (2023) · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I. (2023) · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Clipscore: A reference-free evaluation metric for image captioning
Hessel, J., Holtzman, A., Forbes, M., Bras, R. L., and Choi, Y. (2022) · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., Salimans, T., Ho, J., Fleet, D. J., and Norouzi, M. (2022) · 2022
Cited alongside, same era.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., Schramowski, P., Kundurthy, S., Crowson, K., Schmidt, L., Kaczmarczyk, R., and Jitsev, J. (2022) · 2022
Cited alongside, same era.
Ernie-vilg 2.0: Improving text-to-image diffusion model with knowledge-enhanced mixture-of-denoising-experts
Feng, Z., Zhang, Z., Yu, X., Fang, Y., Li, L., Chen, X., Lu, Y., Liu, J., Yin, W., Feng, S. et al. (2023) · 2023
Cited alongside, same era.
Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf 2 , 8
Betker, J., Goh, G., Jing, L., Brooks, T., Wang, J., Li, L., Ouyang, L., Zhuang, J., Lee, J., Guo, Y. et al. (2023). Improving image generation with better captions · 2023
Cited alongside, same era.
https://photutorial.com/midjourney-statistics/
statistics, M. (2023) · 2023
Cited alongside, same era.
Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models
Qu, Y., Shen, X., He, X., Backes, M., Zannettou, S., and Zhang, Y. (2023) · 2023
Cited alongside, same era.
Later among the works it cites.
Adaptive nonlinear latent transformation for conditional face editing
Huang, Z., Ma, S., Zhang, J., and Shan, H. (2023) · 2023
Later among the works it cites.
Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., Hui, B., Ji, L., Li, M., Lin, J., Lin, R., Liu, D., Liu, G., Lu, C., Lu, K., Ma, J., Men, R., Ren, X., Ren, X., Tan, C., Tan, S., Tu, J., Wang, P., Wang, S., Wang, W., Wu, S., Xu, B., Xu, J., Yang, A., Yang, H., Yang, J., Yang, S., Yao, Y., Yu, B., Yuan, H., Yuan, Z., Zhang, J., Zhang, X., Zhang, Y., Zhang, Z., Zhou, C., Zhou, J., Zhou, X., and Zhu, T. (2023) · 2023
Later among the works it cites.
arXiv preprint arXiv:2303.08774
OpenAI (2023). Gpt-4 technical report · 2023
Later among the works it cites.
Easily accessible text-to-image generation amplifies demographic stereotypes at large scale
Bianchi, F., Kalluri, P., Durmus, E., Ladhak, F., Cheng, M., Nozza, D., Hashimoto, T., Jurafsky, D., Zou, J., and Caliskan, A. (2023) · 2023
Later among the works it cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D. (2023) · 2023
Later among the works it cites.
Social biases through the text-to-image generation lens
Naik, R., and Nushi, B. (2023) · 2023
Later among the works it cites.
arXiv preprint arXiv:2403.05121
Zheng, W., Teng, J., Yang, Z., Wang, W., Chen, J., Gu, X., Dong, Y., Ding, M., and Tang, J. (2024). Cogview3: Finer and faster text-to-image generation via relay diffusion · 2024
Closest in time.
yuzhu-cai/ethical-lens: Ethical-lens
Cai, Y. (2024) · 2024
Closest in time.
Qwen2.5: A party of foundation models
Team, Q. (2024) · 2024
Closest in time.
Dubey, A., Jauhri, A., Pandey, A., and et al. (2024) · 2024
Closest in time.
Yi: Open foundation models by 01.ai
AI, ., :, Young, A., Chen, B., Li, C., Huang, C., Zhang, G., Zhang, G., Li, H., Zhu, J., Chen, J., Chang, J., Yu, K., Liu, P., Liu, Q., Yue, S., Yang, S., Yang, S., Yu, T., Xie, W., Huang, W., Hu, X., Ren, X., Niu, X., Nie, P., Xu, Y., Liu, Y., Wang, Y., Cai, Y., Gu, Z., Liu, Z., and Dai, Z. (2024) · 2024
Closest in time.
mlco2/codecarbon: v2.4.1
Courty, B., Schmidt, V., Luccioni, S., Goyal-Kamal, MarionCoutarel, Feld, B., Lecourt, J., LiamConnell, Saboni, A., Inimaz, supatomic, Léval, M., Blanche, L., Cruveiller, A., ouminasara, Zhao, F., Joshi, A., Bogroff, A., de Lavoreille, H., Laskaris, N., Abati, E., Blank, D., Wang, Z., Catovic, A., Alencon, M., Stęchły, M., Bauer, C., de Araújo, L. O. N., JPW, and MinervaBooks (2024) · 2024
Closest in time.