Fetching the paper…
Reading the bibliography…
We present a novel task and dataset for evaluating the ability of vision and language models to conduct visio-linguistic compositional reasoning, which we call Winoground.
Understanding natural language
Terry Winograd · 1972
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Vicente Ordonez, Girish Kulkarni, and Tamara Berg · 2011
Earlier work this paper cites.
Context models and out-of-context objects
Myung Jin Choi, Antonio Torralba, and Alan S. Willsky · 2012
Earlier work this paper cites.
The winograd schema challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, A. Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehr, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Establishing a human baseline for the winograd schema challenge
David Bender · 2015
Earlier work this paper cites.
Detection of cyberbullying incidents on the instagram social network
Homa Hosseinmardi, Sabrina Arredondo Mattson, Rahat Ibn Rafiq, Richard Han, Qin Lv, and Shivakant Mishra · 2015
Earlier work this paper cites.
Assessing the ability of lstms to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2015
Earlier work this paper cites.
Very deep convolutional networks for largescale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Understanding image and text simultaneously: a dual vision-language machine comprehension task
Nan Ding, Sebastian Goodman, Fei Sha, and Radu Soricut · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Yfcc100m: The new data in multimedia research
Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li · 2016
Earlier work this paper cites.
Content-driven detection of cyberbullying on the instagram social network
Haoti Zhong, Hao Li, Anna Cinzia Squicciarini, Sarah Michele Rajtmajer, Christopher Griffin, David J Miller, and Cornelia Caragea · 2016
Earlier work this paper cites.
Visual7w: Grounded question answering in images
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
Wei-Lun Chao, Hexiang Hu, and Fei Sha · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
First quora dataset release: Question pairs, 2017
Shankar Iyer, Nikhil Dandekar, and Kornel Csernai · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Earlier work this paper cites.
”foil it! find one mismatch between image and language caption”
Ravi Shekhar, Sandro Pezzelle, Yauhen Klimovich, Aurelie Herbelot, Moin Nabi, Enver Sangineto, and Raffaella Bernardi · 2017
Earlier work this paper cites.
A corpus of natural language for visual reasoning
Alane Suhr, Mike Lewis, James Yeh, and Yoav Artzi · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R Bowman · 2017
Earlier work this paper cites.
Vse++: Improving visual-semantic embeddings with hard negatives
Fartash Faghri, David J. Fleet, Jamie Ryan Kiros, and Sanja Fidler · 2018
Earlier work this paper cites.
Colorless green recurrent networks dream hierarchically
Kristina Gulordava, Piotr Bojanowski, Edouard Grave, Tal Linzen, and Marco Baroni · 2018
Cited alongside, same era.
Gender bias in coreference resolution
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Cited alongside, same era.
Do latent tree learning models identify meaningful structure in sentences?
Adina Williams, Andrew Drozdov, and Samuel R. Bowman · 2018
Cited alongside, same era.
Visual entailment task for visually-grounded language learning
Ning Xie, Farley Lai, Derek Doran, and Asim Kadav · 2018
Cited alongside, same era.
Textcaps: a dataset for image captioning with reading comprehension
Oleksii Sidorov, Ronghang Hu, Marcus Rohrbach, and Amanpreet Singh · 2020
Later among the works it cites.
Are we pretraining it right? digging deeper into visio-linguistic pretraining
Amanpreet Singh, Vedanuj Goswami, and Devi Parikh · 2020
Later among the works it cites.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal · 2020
Later among the works it cites.
Investigating novel verb learning in BERT: Selectional preference classes and alternation-based syntactic generalization
Tristan Thrush, Ethan Wilcox, and Roger Levy · 2020
Later among the works it cites.
BLiMP: The benchmark of linguistic minimal pairs for English
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang · 2018
Cited alongside, same era.
A Course in Semantics
Daniel Altshuler, Terence Parsons, and Roger Schwarzschild · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas Dietterich · 2019
Cited alongside, same era.
Evaluating text-to-image matching using binary image selection (bison)
Hexiang Hu, Ishan Misra, and Laurens van der Maaten · 2019
Cited alongside, same era.
Verb argument structure alternations in word and sentence embeddings
Katharina Kann, Alex Warstadt, Adina Williams, and Samuel R. Bowman · 2019
Cited alongside, same era.
Visual semantic reasoning for image-text matching
Kunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li, and Yun Fu · 2019
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2020
Later among the works it cites.
Multimodal datasets: misogyny, pornography, and malignant stereotypes
Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahembwe · 2021
Later among the works it cites.
Automatic generation of contrast sets from scene graphs: Probing the compositional consistency of GQA
Yonatan Bitton, Gabriel Stanovsky, Roy Schwartz, and Michael Elhadad · 2021
Later among the works it cites.
Covr: A test-bed for visually grounded compositional generalization with real images
Ben Bogin, Shivanshu Gupta, Matt Gardner, and Jonathan Berant · 2021
Later among the works it cites.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut · 2021
Later among the works it cites.
Redcaps: Web-curated image-text data created by the people
Karan Desai, Gaurav Kaul, Zubin Aysola, and Justin Johnson · 2021
Later among the works it cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
An empirical study of training end-to-end vision-and-language transformers
Zi-Yi Dou, Yichong Xu, Zhe Gan, Jianfeng Wang, Shuohang Wang, Lijuan Wang, Chenguang Zhu, Zicheng Liu, Michael Zeng, et al · 2021
Later among the works it cites.
Back to square one: Artifact detection, training and commonsense disentanglement in the winograd schema
Yanai Elazar, Hongming Zhang, Yoav Goldberg, and Dan Roth · 2021
Later among the works it cites.
Vision-and-language or vision-for-language? on cross-modal influence in multimodal transformers
Stella Frank, Emanuele Bugliarello, and Desmond Elliott · 2021
Later among the works it cites.
Probing image-language transformers for verb understanding
Lisa Anne Hendricks and Aida Nematzadeh · 2021
Later among the works it cites.
Unit: Multimodal multitask learning with a unified transformer
Ronghang Hu and Amanpreet Singh · 2021
Later among the works it cites.
Vilt: Vision-and-language transformer without convolution or region supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim · 2021
Later among the works it cites.
Hannah Rose Kirk, Bertram Vidgen, Paul Röttger, Tristan Thrush, and Scott A Hale · 2021
Later among the works it cites.
Visually grounded reasoning across languages and cultures
Fangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy, Nigel Collier, and Desmond Elliott · 2021
Later among the works it cites.
Seeing past words: Testing the cross-modal capabilities of pretrained v&l models on counting tasks
Letitia Parcalabescu, Albert Gatt, Anette Frank, and Iacer Calixto · 2021
Later among the works it cites.
Sometimes we want ungrammatical translations
Prasanna Parthasarathi, Koustuv Sinha, Joelle Pineau, and Adina Williams · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Later among the works it cites.
Masked language modeling and the distributional hypothesis: Order word matters pre-training for little
Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, and Douwe Kiela · 2021
Later among the works it cites.
UnNatural Language Inference
Koustuv Sinha, Prasanna Parthasarathi, Joelle Pineau, and Adina Williams · 2021
Later among the works it cites.
Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning
Krishna Srinivasan, Karthik Raman, Jiecao Chen, Michael Bendersky, and Marc Najork · 2021
Later among the works it cites.
Findings of the shared task on troll meme classification in Tamil
Shardul Suryawanshi and Bharathi Raja Chakravarthi · 2021
Later among the works it cites.
Curi: A benchmark for productive concept learning under uncertainty
Ramakrishna Vedantam, Arthur Szlam, Maximillian Nickel, Ari Morcos, and Brenden M Lake · 2021
Later among the works it cites.
Vinvl: Revisiting visual representations in vision-language models
Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao · 2021
Later among the works it cites.
Flava: A foundational language and vision alignment model
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela · 2022
Closest in time.