Fetching the paper…
Reading the bibliography…
Despite the impressive success of text-to-image (TTI) generation models, existing studies overlook the issue of whether these models accurately convey factual information.
Bleu: a Method for Automatic Evaluation of Machine Translation
Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W. 2002 · 2002
Earlier work this paper cites.
On faithfulness and factuality in abstractive summarization
Maynez, J.; Narayan, S.; Bohnet, B.; and McDonald, R. 2020 · 2005
Earlier work this paper cites.
Enhancing students’ learning of factual knowledge
Hew, K. F.; Cheung, W. S.; Hew, K. F.; and Cheung, W. S. 2014 · 2014
Earlier work this paper cites.
CIDEr: Consensus-based image description evaluation
Vedantam, R.; Zitnick, C. L.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
FEVER: a Large-scale Dataset for Fact Extraction and VERification
Thorne, J.; Vlachos, A.; Christodoulopoulos, C.; and Mittal, A. 2018 · 2018
Earlier work this paper cites.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reimers, N.; and Gurevych, I. 2019 · 2019
Earlier work this paper cites.
Language Models are Few-Shot Learners
Brown, T. B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert-Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D. M.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; Sigler, E.; Litwin, M.; Gray, S.; Chess, B.; Clark, J.; Berner, C.; McCandlish, S.; Radford, A.; Sutskever, I.; and Amodei, D. 2020 · 2020
Earlier work this paper cites.
Physics
Urone, P. P.; Hinrichs, R.; Gozuacik, F.; Pattison, D.; and Tabor, C. 2020 · 2020
Earlier work this paper cites.
CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Hessel, J.; Holtzman, A.; Forbes, M.; Bras, R. L.; and Choi, Y. 2021 · 2021
Earlier work this paper cites.
Principles of Biology: An Introduction to Biological Concepts of Biology
O’Grady, E.; Cashmore, J.; Hay, M.; and Wismer, C. 2021 · 2021
Earlier work this paper cites.
Re-imagen: Retrieval-augmented text-to-image generator
Chen, W.; Hu, H.; Saharia, C.; and Cohen, W. W. 2022 · 2022
Earlier work this paper cites.
mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections
Li, C.; Xu, H.; Tian, J.; Wang, W.; Yan, M.; Bi, B.; Ye, J.; Chen, H.; Xu, G.; Cao, Z.; Zhang, J.; Huang, S.; Huang, F.; Zhou, J.; and Si, L. 2022 · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Earlier work this paper cites.
Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E. L.; Ghasemipour, S. K. S.; Lopes, R. G.; Ayan, B. K.; Salimans, T.; Ho, J.; Fleet, D. J.; and Norouzi, M. 2022 · 2022
Cited alongside, same era.
DALL-EVAL: Probing the Reasoning Skills and Social Biases of Text-to-Image Generation Models
Cho, J.; Zala, A.; and Bansal, M. 2023 · 2023
Cited alongside, same era.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Dai, W.; Li, J.; Li, D.; Tiong, A.; Zhao, J.; Wang, W.; Li, B.; Fung, P.; and Hoi, S. 2023 · 2023
Cited alongside, same era.
Breaking Common Sense: WHOOPS! A Vision-and-Language Benchmark of Synthetic and Compositional Images
Guetta, N. B.; Bitton, Y.; Hessel, J.; Schmidt, L.; Elovici, Y.; Stanovsky, G.; and Schwartz, R. 2023 · 2023
Cited alongside, same era.
Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering
Hu, Y.; Liu, B.; Kasai, J.; Wang, Y.; Ostendorf, M.; Krishna, R.; and Smith, N. A. 2023 · 2023
Visual programming for step-by-step text-to-image generation and evaluation
Cho, J.; Zala, A.; and Bansal, M. 2024 · 2024
Closest in time.
The Americans: Student Edition Reconstruction to the 21st Century
Danzer, G. A.; de Alva, J. J. K.; Krieger, L. S.; Wilson, L. E.; and Woloch, N. 2008 · 2024
Closest in time.
Glencoe World History, Student Edition
Jackson J. Spielvogel. 2008 · 2024
Closest in time.
GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation
Li, B.; Lin, Z.; Pathak, D.; Li, J.; Fei, Y.; Wu, K.; Ling, T.; Xia, X.; Zhang, P.; Neubig, G.; et al. 2024 · 2024
Closest in time.
Addressing image hallucination in text-to-image generation through factual image retrieval
Lim, Y.; and Shim, H. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Survey of hallucination in natural language generation
Ji, Z.; Lee, N.; Frieske, R.; Yu, T.; Su, D.; Xu, Y.; Ishii, E.; Bang, Y. J.; Madotto, A.; and Fung, P. 2023 · 2023
Cited alongside, same era.
HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models
Li, J.; Cheng, X.; Zhao, X.; Nie, J.; and Wen, J. 2023a · 2023
Cited alongside, same era.
Evaluating Object Hallucination in Large Vision-Language Models
Li, Y.; Du, Y.; Zhou, K.; Wang, J.; Zhao, W. X.; and Wen, J. 2023b · 2023
Cited alongside, same era.
Factscore: Fine-grained atomic evaluation of factual precision in long form text generation
Min, S.; Krishna, K.; Lyu, X.; Lewis, M.; Yih, W.-t.; Koh, P. W.; Iyyer, M.; Zettlemoyer, L.; and Hajishirzi, H. 2023 · 2023
Cited alongside, same era.
Introduction to Earth Science
Neser, L. 2023 · 2023
Cited alongside, same era.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Podell, D.; English, Z.; Lacey, K.; Blattmann, A.; Dockhorn, T.; Müller, J.; Penna, J.; and Rombach, R. 2023 · 2023
Cited alongside, same era.
Rome: Evaluating pre-trained vision-language models on reasoning beyond visual common sense
Zhou, K.; Lai, E.; Yeong, W. B. A.; Mouratidis, K.; and Jiang, J. 2023 · 2023
Cited alongside, same era.
OpenAI. 2023 · 2024
Closest in time.
GPT-4 Turbo and GPT-4
OpenAI. 2024a · 2024
Closest in time.
GPT-4o mini: advancing cost-efficient intelligence
OpenAI. 2024b · 2024
Closest in time.
Hello GPT-4o
OpenAI. 2024c · 2024
Closest in time.
Google apologizes for ’missing the mark’ after Gemini generated racially diverse Nazis
Robertson, A. 2024 · 2024
Closest in time.
Stable Diffusion v2 Model Card
StabilityAI. 2022 · 2024
Closest in time.
AI-generated images and video are here: how could they shape research?
Wong, C. 2024 · 2024
Closest in time.
What you see is what you read? improving text-image alignment evaluation
Yarom, M.; Bitton, Y.; Changpinyo, S.; Aharoni, R.; Herzig, J.; Lang, O.; Ofek, E.; and Szpektor, I. 2024 · 2024
Closest in time.