Fetching the paper…
Reading the bibliography…
Questions regarding implicitness, ambiguity and underspecification are crucial for understanding the task validity and ethical concerns of multimodal image+text systems, yet have received little attention to date.
Exposing and correcting the gender bias in image captioning datasets and models
Shruti Bhargava and David Forsyth. 2019 · 1912
Earlier work this paper cites.
Collected papers of Charles Sanders Peirce
Charles Hartshorne, Paul Weiss, Arthur W Burks, et al. 1958 · 1958
Earlier work this paper cites.
Closing statement: Linguistics and poetics
Roman Jakobson and Thomas A Sebeok. 1960 · 1960
Earlier work this paper cites.
The determinations of news photographs (1973)
Stuart Hall. 2019 · 1973
Earlier work this paper cites.
Judging people in the news—unconsciously: Effect of camera angle and bodily activity
Lee M Mandell and Donald L Shaw. 1973 · 1973
Earlier work this paper cites.
Basic objects in natural categories
Eleanor Rosch, Carolyn B Mervis, Wayne D Gray, David M Johnson, and Penny Boyes-Braem. 1976 · 1976
Earlier work this paper cites.
Image-music-text
Roland Barthes. 1977 · 1977
Earlier work this paper cites.
Course in General Linguistics
Ferdinand de Saussure. [1916] 1983 · 1983
Earlier work this paper cites.
Analyzing the subject of a picture: a theoretical approach
Sara Shatford. 1986 · 1986
Earlier work this paper cites.
Must scientific diagrams be eliminable?: The case of path analysis
James R Griesemer. 1991 · 1991
Earlier work this paper cites.
Understanding comics: The invisible art
Scott McCloud. 1993 · 1993
Earlier work this paper cites.
Ambiguity, underspecification and discourse interpretation
Massimo Poesio. 1994 · 1994
Earlier work this paper cites.
The construction of social reality
John R Searle. 1995 · 1995
Earlier work this paper cites.
Semantics in generative grammar , volume 1185
Angelika Kratzer and Irene Heim. 1998 · 1998
Earlier work this paper cites.
Producing spoken language
W Levelt. 1999 · 1999
Earlier work this paper cites.
A conceptual framework for indexing visual information at multiple levels
Alejandro Jaimes and Shih-Fu Chang. 2000 · 2000
Earlier work this paper cites.
Web content accessibility guidelines 1.0
Wendy Chisholm, Gregg Vanderheiden, and Ian Jacobs. 2001 · 2001
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2005
Earlier work this paper cites.
Minimal recursion semantics: An introduction
Ann Copestake, Dan Flickinger, Carl Pollard, and Ivan A Sag. 2005 · 2005
Earlier work this paper cites.
Describing images on the web: a survey of current practice and prospects for the future
Helen Petrie, Chandra Harrison, and Sundeep Dev. 2005 · 2005
Earlier work this paper cites.
Some features of "alt" texts associated with images in web pages
Timothy C Craven. 2006 · 2006
Earlier work this paper cites.
Semiotics: the basics
Daniel Chandler. 2007 · 2007
Earlier work this paper cites.
Ways of seeing
John Berger. 2008 · 2008
Earlier work this paper cites.
Comics and sequential art: Principles and practices from the legendary cartoonist
Will Eisner. 2008 · 2008
Earlier work this paper cites.
Semantic underspecification in language processing
Steven Frisson. 2009 · 2009
Earlier work this paper cites.
Between the earth and the air: Multimodality in Arandic sand stories
Jennifer Anne Green. 2009 · 2009
Earlier work this paper cites.
Debunking the myth of the “angry Black woman”: An exploration of anger in young African American women
J Celeste Walley-Jean. 2009 · 2009
Earlier work this paper cites.
Language, technology, and society
Richard Sproat. 2010 · 2010
Earlier work this paper cites.
Computational generation of referring expressions: A survey
Emiel Krahmer and Kees Van Deemter. 2012 · 2012
Earlier work this paper cites.
The Visual Language of Comics: Introduction to the Structure and Cognition of Sequential Images
Neil Cohn. 2013 · 2013
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
Micah Hodosh, Peter Young, and Julia Hockenmaier. 2013 · 2013
Earlier work this paper cites.
Introduction to Peircean visual semiotics
Tony Jappy. 2013 · 2013
Cited alongside, same era.
Orality and literacy
Walter J Ong. 2013 · 2013
Cited alongside, same era.
Generative adversarial nets
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio. 2014 · 2014
Cited alongside, same era.
Microsoft COCO: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Cited alongside, same era.
Predicting entry-level categories
Vicente Ordonez, Wei Liu, Jia Deng, Yejin Choi, Alexander C Berg, and Tamara L Berg. 2015 · 2015
Cited alongside, same era.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015 · 2015
Cited alongside, same era.
Climbing towards NLU: On meaning, form, and understanding in the age of data
Emily M Bender and Alexander Koller. 2020 · 2020
Later among the works it cites.
Experience grounds language
Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, et al. 2020 · 2020
Later among the works it cites.
Capwap: Image captioning with a purpose
Adam Fisch, Kenton Lee, Ming-Wei Chang, Jonathan H Clark, and Regina Barzilay. 2020 · 2020
Later among the works it cites.
The hateful memes challenge: Detecting hate speech in multimodal memes
Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Pratik Ringshia, and Davide Testuggine. 2020 · 2020
Later among the works it cites.
Diversity and inclusion metrics in subset selection
Margaret Mitchell, Dylan Baker, Nyalleng Moorosi, Emily Denton, Ben Hutchinson, Alex Hanna, Timnit Gebru, and Jamie Morgenstern. 2020 · 2020
Later among the works it cites.
Pragmatic issue-sensitive image captioning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
PageSense: Toward stylewise contextual advertising via visual analysis of web pages
Tao Mei, Lusong Li, Xinmei Tian, Dacheng Tao, and Chong-Wah Ngo. 2016 · 2016
Cited alongside, same era.
Seeing through the human reporting bias: Visual classifiers from noisy human-centric labels
Ishan Misra, C Lawrence Zitnick, Margaret Mitchell, and Ross Girshick. 2016 · 2016
Cited alongside, same era.
Stereotyping and bias in the flickr30k dataset
CWJ van Miltenburg. 2016 · 2016
Cited alongside, same era.
Pragmatic factors in image description: The case of negations
Emiel Van Miltenburg, Roser Morante, and Desmond Elliott. 2016 · 2016
Cited alongside, same era.
Physiognomy’s new clothes
Blaise Aguera y Arcas, Margaret Mitchell, and Alexander Todorov. 2017 · 2017
Cited alongside, same era.
Disambiguating visual verbs
Spandana Gella, Frank Keller, and Mirella Lapata. 2017 · 2017
Cited alongside, same era.
Allen Nie, Reuben Cohn-Gordon, and Christopher Potts. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Later among the works it cites.
On the use of human reference data for evaluating automatic image descriptions
Emiel van Miltenburg. 2020 · 2020
Later among the works it cites.
“it’s complicated”: Negotiating accessibility and (mis) representation in image descriptions of race, gender, and disability
Cynthia L Bennett, Cole Gleason, Morgan Klaus Scheuerman, Jeffrey P Bigham, Anhong Guo, and Alexandra To. 2021 · 2021
Later among the works it cites.
Multimodal datasets: misogyny, pornography, and malignant stereotypes
Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahembwe. 2021 · 2021
Later among the works it cites.
CogView: Mastering text-to-image generation via transformers
Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, and Jie Tang. 2021 · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021 · 2021
Later among the works it cites.
Concadia: Tackling image accessibility with context
Elisa Kreiss, Noah D Goodman, and Christopher Potts. 2021 · 2021
Later among the works it cites.
GLIDE: Towards photorealistic image generation and editing with text-guided diffusion models
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021 · 2021
Later among the works it cites.
AI and the everything in the whole wide world benchmark
Inioluwa Deborah Raji, Emily Denton, Emily M Bender, Alex Hanna, and Amandalynne Paullada. 2021 · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021 · 2021
Later among the works it cites.
Targeting the benchmark: On methodology in current natural language processing research
David Schlangen. 2021 · 2021
Later among the works it cites.
WIT: Wikipedia-based image text dataset for multimodal multilingual machine learning
Krishna Srinivasan, Karthik Raman, Jiecao Chen, Michael Bendersky, and Marc Najork. 2021 · 2021
Later among the works it cites.
Cross-modal contrastive learning for text-to-image generation
Han Zhang, Jing Yu Koh, Jason Baldridge, Honglak Lee, and Yinfei Yang. 2021 · 2021
Later among the works it cites.
Pali: A jointly-scaled multilingual language-image model
Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, et al. 2022 · 2022
Closest in time.
DALL-Eval: Probing the reasoning skills and social biases of text-to-image generative transformers
Jaemin Cho, Abhay Zala, and Mohit Bansal. 2022 · 2022
Closest in time.
Underspecification presents challenges for credibility in modern machine learning
Alexander D’Amour, Katherine Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Alex Beutel, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew D Hoffman, et al. 2022 · 2022
Closest in time.
Handling and presenting harmful text
Leon Derczynski, Hannah Rose Kirk, Abeba Birhane, and Bertie Vidgen. 2022 · 2022
Closest in time.
Make-a-scene: Scene-based text-to-image generation with human priors
Oran Gafni, Adam Polyak, Oron Ashual, Shelly Sheynin, Devi Parikh, and Yaniv Taigman. 2022 · 2022
Closest in time.
What’s in an alt tag? exploring caption content priorities through collaborative captioning
Annika Muehlbradt and Shaun K Kane. 2022 · 2022
Closest in time.
Justice in misinformation detection systems: An analysis of algorithms, stakeholders, and potential harms
Terrence Neumann, Maria De-Arteaga, and Sina Fazelpour. 2022 · 2022
Closest in time.
Hierarchical text-conditional image generation with CLIP latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022 · 2022
Closest in time.
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al. 2022 · 2022
Closest in time.
Scaling autoregressive models for content-rich text-to-image generation
Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, et al. 2022 · 2022
Closest in time.