Fetching the paper…
Reading the bibliography…
AI-based text-to-image models do not only excel at generating realistic images, they also give designers more and more fine-grained control over the image content.
J. Cho, J. Lei, H. Tan, and M. Bansal, “Unifying vision-and-language tasks via text generation,” in Proceedings of the 38th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, vol. 139. PMLR, 18–24 Jul 2021, pp. 1931–1942
1942
Earlier work this paper cites.
G. A. Miller, “Wordnet: a lexical database for english,” Commun. ACM , vol. 38, no. 11, p. 39–41, Nov. 1995
1995
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th Annual Meeting on Association for Computational Linguistics , ser. ACL ’02. USA: Association for Computational Linguistics, 2002, p. 311–318
2002
Earlier work this paper cites.
P. Jenkins, A. Farag, S. Wang, and Z. Li, “Unsupervised representation learning of spatial data via multimodal embedding,” in Proceedings of the 28th ACM international conference on information and knowledge management , 2019, pp. 1993–2002
2002
Earlier work this paper cites.
C.-Y. Lin, “ROUGE: A package for automatic evaluation of summaries,” in Text Summarization Branches Out . Barcelona, Spain: Association for Computational Linguistics, Jul. 2004, pp. 74–81
2004
Earlier work this paper cites.
S. Banerjee and A. Lavie, “METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,” in Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization . Ann Arbor, Michigan: Association for Computational Linguistics, Jun. 2005, pp. 65–72
2005
Earlier work this paper cites.
A. Gretton, K. Borgwardt, M. Rasch, B. Schölkopf, and A. Smola, “A kernel method for the two-sample-problem,” Advances in neural information processing systems , vol. 19, 2006
2006
Earlier work this paper cites.
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman, “The PASCAL Visual Object Classes Challenge 2008 (VOC2008) Results,” http://www.pascal-network.org/challenges/VOC/voc2008/workshop/index.html
2008
Earlier work this paper cites.
C. Rashtchian, P. Young, M. Hodosh, and J. Hockenmaier, “Collecting image annotations using amazon’s mechanical turk,” in Proceedings of the NAACL HLT 2010 workshop on creating speech and language data with Amazon’s Mechanical Turk , 2010, pp. 139–147
2010
Earlier work this paper cites.
V. Ordonez, G. Kulkarni, and T. Berg, “Im2text: Describing images using 1 million captioned photographs,” in Advances in Neural Information Processing Systems , vol. 24. Curran Associates, Inc., 2011
2011
Earlier work this paper cites.
M. Hodosh, P. Young, and J. Hockenmaier, “Framing image description as a ranking task: Data, models and evaluation metrics,” Journal of Artificial Intelligence Research , vol. 47, pp. 853–899, 2013
2013
Earlier work this paper cites.
C. L. Zitnick and D. Parikh, “Bringing semantics into focus using visual abstraction,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , Jun. 2013
2013
Earlier work this paper cites.
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier, “From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,” Transactions of the Association for Computational Linguistics , vol. 2, pp. 67–78, 2014
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13 . Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
R. Vedantam, C. Lawrence Zitnick, and D. Parikh, “Cider: Consensus-based image description evaluation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , Jun. 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik, “Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,” in 2015 IEEE International Conference on Computer Vision (ICCV) , 2015, pp. 2641–2649
2015
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “Vqa: Visual question answering,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV) , Dec. 2015
2015
Earlier work this paper cites.
T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, X. Chen, and X. Chen, “Improved techniques for training gans,” in Advances in Neural Information Processing Systems , vol. 29. Curran Associates, Inc., 2016
2016
Earlier work this paper cites.
P. Anderson, B. Fernando, M. Johnson, and S. Gould, “Spice: Semantic propositional image caption evaluation,” in Computer Vision – ECCV 2016 . Cham: Springer International Publishing, 2016, pp. 382–398
2016
Earlier work this paper cites.
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 2818–2826
2016
Earlier work this paper cites.
D. Lopez-Paz and M. Oquab, “Revisiting classifier two-sample tests,” in International Conference on Learning Representations , 2016
2016
Earlier work this paper cites.
T. Che, Y. Li, A. Jacob, Y. Bengio, and W. Li, “Mode regularized generative adversarial networks,” in International Conference on Learning Representations , 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
J. Mao, J. Xu, Y. Jing, and A. Yuille, “Training and evaluating multimodal word embeddings with large-scale web annotated images,” in Proceedings of the 30th International Conference on Neural Information Processing Systems , ser. NIPS’16. Red Hook, NY, USA: Curran Associates Inc., 2016, p. 442–450
2016
Earlier work this paper cites.
A. Rohrbach, A. Torabi, M. Rohrbach, N. Tandon, C. Pal, H. Larochelle, A. Courville, and B. Schiele, “Movie description,” 2016
2016
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, M. S. Bernstein, and L. Fei-Fei, “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” International Journal of Computer Vision , vol. 123, pp. 32–73, 2016
2016
Earlier work this paper cites.
X. Wu, K. Xu, and P. Hall, “A survey of image synthesis and editing with generative adversarial networks,” Tsinghua Science and Technology , vol. 22, no. 6, pp. 660–674, 2017
2017
Earlier work this paper cites.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” Advances in neural information processing systems , vol. 30, p. 6629–6640, 2017
2017
Earlier work this paper cites.
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering,” in Conference on Computer Vision and Pattern Recognition (CVPR) , Jul. 2017
2017
Earlier work this paper cites.
Z. Wang, C. Lu, Y. Wang, F. Bao, C. Li, H. Su, and J. Zhu, “Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distillation,” in Thirty-seventh Conference on Neural Information Processing Systems , 2023
2017
Earlier work this paper cites.
Y. Cui, G. Yang, A. Veit, X. Huang, and S. Belongie, “Learning to evaluate image captioning,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , Jun. 2018
2018
Earlier work this paper cites.
T. Xu, P. Zhang, Q. Huang, H. Zhang, Z. Gan, X. Huang, and X. He, “Attngan: Fine-grained text to image generation with attentional generative adversarial networks,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Los Alamitos, CA, USA: IEEE Computer Society, Jun. 2018, pp. 1316–1324
2018
Earlier work this paper cites.
M. Bińkowski, D. J. Sutherland, M. Arbel, and A. Gretton, “Demystifying mmd gans,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
M. S. Sajjadi, O. Bachem, M. Lucic, O. Bousquet, and S. Gelly, “Assessing generative models via precision and recall,” Advances in neural information processing systems , vol. 31, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
S. Barratt and R. Sharma, “A note on the inception score,” arXiv preprint arXiv:1801.01973 , 2018
2018
Earlier work this paper cites.
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2018, pp. 2556–2565
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Kiros, W. Chan, and G. Hinton, “Illustrative language understanding: Large-scale visual grounding with image search,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2018, pp. 922–933
2018
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
P. Madhyastha, J. Wang, and L. Specia, “VIFIDEL: Evaluating the visual fidelity of image descriptions,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . Florence, Italy: Association for Computational Linguistics, Jul. 2019, pp. 6539–6550
2019
Earlier work this paper cites.
S. Ravuri and O. Vinyals, “Classification accuracy score for conditional generative models,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
T. Kynkäänniemi, T. Karras, S. Laine, J. Lehtinen, and T. Aila, “Improved precision and recall metric for assessing generative models,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
T. Karras, S. Laine, and T. Aila, “A style-based generator architecture for generative adversarial networks,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , Jun. 2019, pp. 4401–4410
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
K. Li, Y. Zhang, K. Li, Y. Li, and Y. Fu, “Visual semantic reasoning for image-text matching,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , Oct. 2019
2019
Earlier work this paper cites.
J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” in Advances in Neural Information Processing Systems , vol. 32. Curran Associates, Inc., 2019
2019
Earlier work this paper cites.
H. Agrawal, K. Desai, Y. Wang, X. Chen, R. Jain, M. Johnson, D. Batra, D. Parikh, S. Lee, and P. Anderson, “nocaps: novel object captioning at scale,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , Oct. 2019
2019
Cited alongside, same era.
R. Zellers, Y. Bisk, A. Farhadi, and Y. Choi, “From recognition to cognition: Visual commonsense reasoning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , Jun. 2019
2019
Cited alongside, same era.
N. Xie, F. Lai, D. Doran, and A. Kadav, “Visual entailment task for visually-grounded language learning,” 2019
2019
Cited alongside, same era.
H. Lee, S. Yoon, F. Dernoncourt, D. S. Kim, T. Bui, and K. Jung, “ViLBERTScore: Evaluating image caption using vision-and-language BERT,” in Proceedings of the First Workshop on Evaluation and Comparison of NLP Systems . Online: Association for Computational Linguistics, Nov. 2020, pp. 34–39
2020
Cited alongside, same era.
C. Zhang, C. Zhang, M. Zhang, and I. S. Kweon, “Text-to-image diffusion models in generative ai: A survey,” 2023
2023
Later among the works it cites.
F.-A. Croitoru, V. Hondru, R. T. Ionescu, and M. Shah, “Diffusion models in vision: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 9, pp. 10 850–10 869, 2023
2023
Later among the works it cites.
R. Po, W. Yifan, V. Golyanik, K. Aberman, J. T. Barron, A. H. Bermano, E. R. Chan, T. Dekel, A. Holynski, A. Kanazawa, C. K. Liu, L. Liu, B. Mildenhall, M. Nießner, B. Ommer, C. Theobalt, P. Wonka, and G. Wetzstein, “State of the art on diffusion models for visual computing,” 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Hinz, S. Heinrich, and S. Wermter, “Semantic object accuracy for generative text-to-image synthesis,” IEEE transactions on pattern analysis and machine intelligence , vol. 44, no. 3, pp. 1552–1565, 2020
2020
Cited alongside, same era.
S. Gu, J. Bao, D. Chen, and F. Wen, “Giqa: Generated image quality assessment,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XI 16 . Springer, 2020, pp. 369–385
2020
Cited alongside, same era.
Y.-C. Chen, L. Li, L. Yu, A. El Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu, “Uniter: Universal image-text representation learning,” in European conference on computer vision . Springer, 2020, pp. 104–120
2020
Cited alongside, same era.
Z. Gan, Y.-C. Chen, L. Li, C. Zhu, Y. Cheng, and J. Liu, “Large-scale adversarial training for vision-and-language representation learning,” Advances in Neural Information Processing Systems , vol. 33, pp. 6616–6628, 2020
2020
Cited alongside, same era.
J. Liang, W. Pei, and F. Lu, “Cpgan: Content-parsing generative adversarial networks for text-to-image synthesis,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part IV 16 . Springer, 2020, pp. 491–508
2020
Cited alongside, same era.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
T. Zhang*, V. Kishore*, F. Wu*, K. Q. Weinberger, and Y. Artzi, “Bertscore: Evaluating text generation with bert,” in International Conference on Learning Representations , 2020
2020
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
L. Zhang, Z. Xu, C. Barnes, Y. Zhou, Q. Liu, H. Zhang, S. Amirghodsi, Z. Lin, E. Shechtman, and J. Shi, “Perceptual artifacts localization for image synthesis tasks,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 7579–7590
2023
Later among the works it cites.
B. Gordon, Y. Bitton, Y. Shafir, R. Garg, X. Chen, D. Lischinski, D. Cohen-Or, and I. Szpektor, “Mismatch quest: Visual and textual feedback for image-text misalignment,” 2023
2023
Later among the works it cites.
J. Li, D. Li, S. Savarese, and S. C. H. Hoi, “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in International Conference on Machine Learning , 2023
2023
Later among the works it cites.
R. OpenAI, “Gpt-4 technical report. arxiv 2303.08774,” View in Article , vol. 2, p. 13, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
K. Huang, K. Sun, E. Xie, Z. Li, and X. Liu, “T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation,” arXiv preprint arXiv: 2307.06350 , 2023
2023
Later among the works it cites.
F. Betti, J. Staiano, L. Baraldi, L. Baraldi, R. Cucchiara, and N. Sebe, “Let’s vice! mimicking human cognitive behavior in image generation evaluation,” in Proceedings of the 31st ACM International Conference on Multimedia , 2023, pp. 9306–9312
2023
Later among the works it cites.
Y. Hu, B. Liu, J. Kasai, Y. Wang, M. Ostendorf, R. Krishna, and N. A. Smith, “Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , Oct. 2023, pp. 20 406–20 417
2023
Later among the works it cites.
M. Yarom, Y. Bitton, S. Changpinyo, R. Aharoni, J. Herzig, O. Lang, E. Ofek, and I. Szpektor, “What you see is what you read? improving text-image alignment evaluation,” 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
H. Chefer, Y. Alaluf, Y. Vinker, L. Wolf, and D. Cohen-Or, “Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models,” ACM Transactions on Graphics (TOG) , vol. 42, no. 4, pp. 1–10, 2023
2023
Later among the works it cites.
Y. Lu, X. Yang, X. Li, X. E. Wang, and W. Y. Wang, “LLMScore: Unveiling the power of large language models in text-to-image synthesis evaluation,” in Thirty-seventh Conference on Neural Information Processing Systems , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
N. Ruiz, Y. Li, V. Jampani, Y. Pritch, M. Rubinstein, and K. Aberman, “Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , Jun. 2023, pp. 22 500–22 510
2023
Later among the works it cites.
J. Wang, K. C. Chan, and C. C. Loy, “Exploring clip for assessing the look and feel of images,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, no. 2, 2023, pp. 2555–2563
2023
Later among the works it cites.
J. Cho, A. Zala, and M. Bansal, “Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , Oct. 2023, pp. 3043–3054
2023
Later among the works it cites.
W. Feng, X. He, T.-J. Fu, V. Jampani, A. R. Akula, P. Narayana, S. Basu, X. E. Wang, and W. Y. Wang, “Training-free structured diffusion guidance for compositional text-to-image synthesis,” in The Eleventh International Conference on Learning Representations , 2023
2023
Later among the works it cites.
P. Schramowski, M. Brack, B. Deiseroth, and K. Kersting, “Safe latent diffusion: Mitigating inappropriate degeneration in diffusion models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , Jun. 2023, pp. 22 522–22 531
2023
Later among the works it cites.
T. Lee, M. Yasunaga, C. Meng, Y. Mai, J. S. Park, A. Gupta, Y. Zhang, D. Narayanan, H. B. Teufel, M. Bellagente, M. Kang, T. Park, J. Leskovec, J.-Y. Zhu, L. Fei-Fei, J. Wu, S. Ermon, and P. Liang, “Holistic evaluation of text-to-image models,” in Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track , 2023
2023
Later among the works it cites.
S. Wang, C. Saharia, C. Montgomery, J. Pont-Tuset, S. Noy, S. Pellegrini, Y. Onoe, S. Laszlo, D. J. Fleet, R. Soricut, J. Baldridge, M. Norouzi, P. Anderson, and W. Chan, “Imagen editor and editbench: Advancing and evaluating text-guided image inpainting,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , Jun. 2023, pp. 18 359–18 369
2023
Later among the works it cites.
J. Urbanek, F. Bordes, P. Astolfi, M. Williamson, V. Sharma, and A. Romero-Soriano, “A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions,” 2023
2023
Later among the works it cites.
Z. Ma, J. Hong, M. O. Gul, M. Gandhi, I. Gao, and R. Krishna, “Crepe: Can vision-language foundation models reason compositionally?” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , Jun. 2023, pp. 10 910–10 921
2023
Later among the works it cites.
C.-Y. Hsieh, J. Zhang, Z. Ma, A. Kembhavi, and R. Krishna, “Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality,” in Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track , 2023
2023
Later among the works it cites.
N. Dehouche and K. Dehouche, “What’s in a text-to-image prompt? the potential of stable diffusion in visual arts education,” Heliyon , vol. 9, no. 6, 2023
2023
Later among the works it cites.
L. Yang, Z. Zhang, Y. Song, S. Hong, R. Xu, Y. Zhao, W. Zhang, B. Cui, and M.-H. Yang, “Diffusion models: A comprehensive survey of methods and applications,” ACM Computing Surveys , vol. 56, no. 4, pp. 1–39, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
C.-H. Lin, J. Gao, L. Tang, T. Takikawa, X. Zeng, X. Huang, K. Kreis, S. Fidler, M.-Y. Liu, and T.-Y. Lin, “Magic3d: High-resolution text-to-3d content creation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , Jun. 2023, pp. 300–309
2023
Later among the works it cites.
G. Metzer, E. Richardson, O. Patashnik, R. Giryes, and D. Cohen-Or, “Latent-nerf for shape-guided generation of 3d shapes and textures,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , Jun. 2023, pp. 12 663–12 673
2023
Later among the works it cites.
R. Chen, Y. Chen, N. Jiao, and K. Jia, “Fantasia3d: Disentangling geometry and appearance for high-quality text-to-3d content creation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , Oct. 2023, pp. 22 246–22 256
2023
Later among the works it cites.
C. Xu, Y. Xu, H. Zhang, X. Xu, and S. He, “Dreamanime: Learning style-identity textual disentanglement for anime and beyond,” IEEE Transactions on Visualization and Computer Graphics , 2024
2024
Closest in time.
J. Xing, M. Xia, Y. Liu, Y. Zhang, Y. He, H. Liu, H. Chen, X. Cun, X. Wang, Y. Shan et al. , “Make-your-video: Customized video generation using textual and structural guidance.” IEEE Transactions on Visualization and Computer Graphics , 2024
2024
Closest in time.
2024
Closest in time.
P. Grimal, H. Le Borgne, O. Ferret, and J. Tourille, “Tiam - a metric for evaluating alignment in text-to-image generation,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , Jan. 2024, pp. 2890–2899
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Lin, D. Pathak, B. Li, J. Li, X. Xia, G. Neubig, P. Zhang, and D. Ramanan, “Evaluating text-to-visual generation with image-to-text generation,” 2024
2024
Closest in time.
J. Wang, H. Duan, G. Zhai, and X. Min, “Understanding and evaluating human preferences for ai generated images with instruction tuning,” 2024
2024
Closest in time.
A. Ray, F. Radenovic, A. Dubey, B. Plummer, R. Krishna, and K. Saenko, “cola: A benchmark for compositional text-to-image retrieval,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
X. Wu, K. Sun, F. Zhu, R. Zhao, and H. Li, “Human preference score: Better aligning text-to-image models with human preference,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , Oct. 2023, pp. 2096–2105
2096
Closest in time.