2022

DALLE-2 is Seeing Double: Flaws in Word-to-Concept Mapping in Text2Image Models

Rassin, Royi, Ravfogel, Shauli, Goldberg, Yoav

Understand

We study the way DALLE-2 maps symbols (words) in the prompt to their references (entities or properties of entities in the generated image).

  • We show that in stark contrast to the way human process language, DALLE-2 does not follow the constraint that each word has a single role in the interpretation, and sometimes re-use the same symbol for different purposes.
  • We collect a set of stimuli that reflect the phenomenon: we show that DALLE-2 depicts both senses of nouns with multiple senses at once; and that a given word can modify the properties of two distinct entities in the image, or can be depicted as one object and also modify the properties of another object, creating a semantic leakage of properties between entities.
  • Taken together, our study highlights the differences between DALLE-2 and human language processing and opens an avenue for future study on the inductive biases of text-to-image models.

Built on

  • Toward a unified theory of the multitasking continuum: From concurrent performance to task switching, interruption, and resumption

    Dario D Salvucci, Niels A Taatgen, and Jelmer P Borst. 2009 · 2009

    Earlier work this paper cites.

  • Dall·e mini

    Boris Dayma, Suraj Patil, Pedro Cuenca, Khalid Saifullah, Tanishq Abraham, Phúc Lê Khac, Luke Melas, and Ritobrata Ghosh. 2021 · 2021

    Earlier work this paper cites.

  • Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021

    Earlier work this paper cites.

  • High-resolution image synthesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2021 · 2021

    Earlier work this paper cites.

Similar

  • Testing relational understanding in text-guided image generation

    Colin Conwell and Tomer Ullman. 2022 · 2022

    Cited alongside, same era.

  • Underspecification in scene description-to-depiction tasks

    Original

    Ben Hutchinson, Jason Baldridge, and Vinodkumar Prabhakaran. 2022 · 2022

    Cited alongside, same era.

  • A very preliminary analysis of dall-e 2

    Gary Marcus, Ernest Davis, and Scott Aaronson. 2022 · 2022

    Cited alongside, same era.

  • Hierarchical text-conditional image generation with clip latents

    Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022 · 2022

    Cited alongside, same era.

Then

  • Photorealistic text-to-image diffusion models with deep language understanding

    Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi. 2022 · 2022

    Closest in time.

  • Scaling autoregressive models for content-rich text-to-image generation

    Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, Ben Hutchinson, Wei Han, Zarana Parekh, Xin Li, Han Zhang, Jason Baldridge, and Yonghui Wu. 2022 · 2022

    Closest in time.

  • Training-free structured diffusion guidance for compositional text-to-image synthesis

    Anonymous. 2023 · 2023

    Closest in time.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…