Fetching the paper…
Reading the bibliography…
The ability to integrate context, including perceptual and temporal cues, plays a pivotal role in grounding the meaning of a linguistic utterance.
Binary Image Selection (BISON): Interpretable Evaluation of Visual Grounding
Hexiang Hu, Ishan Misra, and Laurens van der Maaten. 2019 · 1901
Earlier work this paper cites.
Visual entailment: A novel task for fine-grained image understanding
Ning Xie, Farley Lai, Derek Doran, and Asim Kadav. 2019 · 1901
Earlier work this paper cites.
Meaning." the philosophical review 66: 377-88. 1969
H Paul Grice. 1957 · 1969
Earlier work this paper cites.
Pragmatics and time
Deirdre Wilson and Dan Sperber. 1998 · 1998
Earlier work this paper cites.
Language, thought and compositionality
Jerry Fodor. 2001 · 2001
Earlier work this paper cites.
Grounding cognition: The role of perception and action in memory, language, and thinking
Diane Pecher and Rolf A Zwaan. 2005 · 2005
Earlier work this paper cites.
Converting video formats with ffmpeg
Suramya Tomar. 2006 · 2006
Earlier work this paper cites.
An overview of the Tesseract OCR engine
Ray Smith. 2007 · 2007
Earlier work this paper cites.
No-reference image quality assessment in the spatial domain
Anish Mittal, Anush Krishna Moorthy, and Alan Conrad Bovik. 2012 · 2012
Earlier work this paper cites.
A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching
Pradipto Das, Chenliang Xu, Richard F Doell, and Jason J Corso. 2013 · 2013
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg. 2014 · 2014
Earlier work this paper cites.
Reasoning about Pragmatics with Neural Listeners and Speakers
Jacob Andreas and Dan Klein. 2016 · 2016
Earlier work this paper cites.
Pragmatic language interpretation as probabilistic inference
Noah D Goodman and Michael C Frank. 2016 · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan Yuille, and Kevin Murphy. 2016 · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016 · 2016
Cited alongside, same era.
Visual dialog
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José MF Moura, Devi Parikh, and Dhruv Batra. 2017 · 2017
Cited alongside, same era.
Guesswhat?! visual object discovery through multi-modal dialogue
Harm de Vries, Florian Strub, Sarath Chandar, Olivier Pietquin, Hugo Larochelle, and Aaron C. Courville. 2017 · 2017
Cited alongside, same era.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017 · 2017
Cited alongside, same era.
spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing
Matthew Honnibal and Ines Montani. 2017 · 2017
Cited alongside, same era.
Visual question answering on image sets
Ankan Bansal, Yuting Zhang, and Rama Chellappa. 2020 · 2020
Later among the works it cites.
UNITER: UNiversal Image-TExt Representation Learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020 · 2020
Later among the works it cites.
Negated and misprimed probes for pretrained language models: Birds can talk, but cannot fly
Nora Kassner and Hinrich Schütze. 2020 · 2020
Later among the works it cites.
The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, Tom Duerig, and Vittorio Ferrari. 2020 · 2020
Later among the works it cites.
COVR: A test-bed for visually grounded compositional generalization with real images
Ben Bogin, Shivanshu Gupta, Matt Gardner, and Jonathan Berant. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ramakrishna Vedantam, Samy Bengio, Kevin Murphy, Devi Parikh, and Gal Chechik. 2017 · 2017
Cited alongside, same era.
Pragmatically informative image captioning with character-level inference
Reuben Cohn-Gordon, Noah Goodman, and Christopher Potts. 2018 · 2018
Cited alongside, same era.
Learning to describe differences between pairs of similar images
Harsh Jhamtani and Taylor Berg-Kirkpatrick. 2018b · 2018
Cited alongside, same era.
Neural Naturalist: Generating Fine-Grained Image Comparisons
Maxwell Forbes, Christine Kaeser-Chen, Piyush Sharma, and Serge Belongie. 2019 · 2019
Cited alongside, same era.
Are we modeling the task or the annotator? an investigation of annotator bias in natural language understanding datasets
Mor Geva, Yoav Goldberg, and Jonathan Berant. 2019 · 2019
Cited alongside, same era.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning. 2019 · 2019
Cited alongside, same era.
Video storytelling: Textual summaries for events
Junnan Li, Yongkang Wong, Qi Zhao, and Mohan S Kankanhalli. 2019 · 2019
Cited alongside, same era.
Multimodal Pretraining Unmasked: A Meta-Analysis and a Unified Framework of Vision-and-Language BERTs
Emanuele Bugliarello, Ryan Cotterell, Naoaki Okazaki, and Desmond Elliott. 2021 · 2021
Later among the works it cites.
Probing Image-Language Transformers for Verb Understanding
Lisa Anne Hendricks and Aida Nematzadeh. 2021 · 2021
Later among the works it cites.
Understanding by understanding not: Modeling negation in language models
Arian Hosseini, Siva Reddy, Dzmitry Bahdanau, R Devon Hjelm, Alessandro Sordoni, and Aaron Courville. 2021 · 2021
Later among the works it cites.
Image change captioning by learning from an auxiliary task
Mehrdad Hosseinzadeh and Yang Wang. 2021 · 2021
Later among the works it cites.
Visually grounded reasoning across languages and cultures
Fangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy, Nigel Collier, and Desmond Elliott. 2021 · 2021
Later among the works it cites.
Thinking Fast and Slow: Efficient Text-to-Visual Retrieval with Transformers
Antoine Miech, Jean-Baptiste Alayrac, Ivan Laptev, Josef Sivic, and Andrew Zisserman. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Later among the works it cites.
A list of color, emotion, and human body part concepts
Annika Tjuka. 2021 · 2021
Later among the works it cites.
L2C: Describing visual differences needs semantic understanding of individuals
An Yan, Xin Wang, Tsu-Jui Fu, and William Yang Wang. 2021 · 2021
Later among the works it cites.