Fetching the paper…
Reading the bibliography…
We propose VALSE (Vision And Language Structured Evaluation), a novel benchmark designed for testing general-purpose pretrained vision and language (V&L) models for their visio-linguistic grounding capabilities on specific linguistic phenomena.
Visual Entailment: A Novel Task for Fine-Grained Image Understanding
Ning Xie, Farley Lai, Derek Doran, and Asim Kadav. 2019 · 1901
Earlier work this paper cites.
WordNet: An Electronic Lexical Database
Christiane Fellbaum. 1998 · 1998
Earlier work this paper cites.
Behind the scene: Revealing the secrets of pre-trained vision-and-language models
Jize Cao, Zhe Gan, Yu Cheng, Licheng Yu, Yen-Chun Chen, and Jingjing Liu. 2020 · 2005
Earlier work this paper cites.
SimpleNLG: A realisation engine for practical applications
Albert Gatt and Ehud Reiter. 2009 · 2009
Earlier work this paper cites.
Language models are open knowledge graphs
Chenguang Wang, Xiao Liu, and Dawn Song. 2020a · 2010
Earlier work this paper cites.
The winograd schema challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern. 2012 · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Microsoft COCO Captions : Data Collection and Evaluation Server
Xinlei Chen, Hao Fang, Tsung-yi Lin, Ramakrishna Vedantam, C Lawrence Zitnick, Saurabh Gupta, and Piotr Doll. 2015 · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. 2015 · 2015
Earlier work this paper cites.
Situation recognition: Visual semantic role labeling for image understanding
Mark Yatskar, Luke Zettlemoyer, and Ali Farhadi. 2016 · 2016
Earlier work this paper cites.
Visual7W: Grounded Question Answering in Images
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei. 2016 · 2016
Earlier work this paper cites.
Visual Dialog
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José M.F. Moura, Devi Parikh, and Dhruv Batra. 2017 · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017 · 2017
Earlier work this paper cites.
Vision and language integration: Moving beyond objects
Ravi Shekhar, Sandro Pezzelle, Aurélie Herbelot, Moin Nabi, Enver Sangineto, and Raffaella Bernardi. 2017a · 2017
Earlier work this paper cites.
Visual referring expression recognition: What do systems actually learn?
Volkan Cirik, Louis-Philippe Morency, and Taylor Berg-Kirkpatrick. 2018 · 2018
Earlier work this paper cites.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018 · 2018
Earlier work this paper cites.
Defoiling foiled image captions
Pranava Swaroop Madhyastha, Josiah Wang, and Lucia Specia. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Counterfactual fairness in text classification through robustness
Sahaj Garg, Vincent Perot, Nicole Limtiaco, Ankur Taly, Ed H. Chi, and Alex Beutel. 2019 · 2019
Cited alongside, same era.
Certified robustness to adversarial word substitutions
Robin Jia, Aditi Raghunathan, Kerem Göksel, and Percy Liang. 2019 · 2019
Cited alongside, same era.
Challenges and prospects in vision and language research
Kushal Kafle, Robik Shrestha, and Christopher Kanan. 2019 · 2019
Cited alongside, same era.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019 · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Cited alongside, same era.
VIFIDEL: Evaluating the visual fidelity of image descriptions
Adversarial NLI: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020 · 2020
Later among the works it cites.
Grounded situation recognition
Sarah Pratt, Mark Yatskar, Luca Weihs, Ali Farhadi, and Aniruddha Kembhavi. 2020 · 2020
Later among the works it cites.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Later among the works it cites.
Vl-bert: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020 · 2020
Later among the works it cites.
Interpreting predictions of NLP models
Eric Wallace, Matt Gardner, and Sameer Singh. 2020 · 2020
Later among the works it cites.
Gradient-based analysis of NLP models is manipulable
Junlin Wang, Jens Tuyls, Eric Wallace, and Sameer Singh. 2020b · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pranava Madhyastha, Josiah Wang, and Lucia Specia. 2019 · 2019
Cited alongside, same era.
Combining fact extraction and verification with neural semantic matching networks
Yixin Nie, Haonan Chen, and Mohit Bansal. 2019 · 2019
Cited alongside, same era.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Beyond task success: A closer look at jointly learning to see, ask, and GuessWhat
Ravi Shekhar, Aashish Venkatesh, Tim Baumgärtner, Elia Bruni, Barbara Plank, Raffaella Bernardi, and Raquel Fernández. 2019b · 2019
Cited alongside, same era.
A corpus for reasoning about natural language grounded in photographs
Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi. 2019 · 2019
Cited alongside, same era.
LXMERT: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Julien Chaumond, Lysandre Debut, Victor Sanh, Clement Delangue, Anthony Moi, Pierric Cistac, Morgan Funtowicz, Joe Davison, Sam Shleifer, et al. 2020 · 2020
Later among the works it cites.
GRUEN for evaluating linguistic quality of generated text
Wanzheng Zhu and Suma Bhat. 2020 · 2020
Later among the works it cites.
Linguistic issues behind visual question answering
Raffaella Bernardi and Sandro Pezzelle. 2021 · 2021
Closest in time.
Automatic generation of contrast sets from scene graphs: Probing the compositional consistency of GQA
Yonatan Bitton, Gabriel Stanovsky, Roy Schwartz, and Michael Elhadad. 2021 · 2021
Closest in time.
Multimodal neurons in artificial neural networks
Gabriel Goh, Nick Cammarata †, Chelsea Voss †, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah. 2021 · 2021
Closest in time.
Interpretable visual reasoning: A survey
Feijuan He, Yaxian Wang, Xianglin Miao, and Xia Sun. 2021 · 2021
Closest in time.
Probing Image-Language Transformers for Verb Understanding
Lisa Anne Hendricks and Aida Nematzadeh. 2021 · 2021
Closest in time.
Vilt: Vision-and-language transformer without convolution or region supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021 · 2021
Closest in time.
Value: A multi-task benchmark for video-and-language understanding evaluation
Linjie Li, Jie Lei, Zhe Gan, Licheng Yu, Yen-Chun Chen, Rohit Pillai, Yu Cheng, Luowei Zhou, Xin Eric Wang, William Yang Wang, et al. 2021 · 2021
Closest in time.
Unicorn on rainbow: A universal commonsense reasoning model on a new multitask benchmark
Nicholas Lourie, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021 · 2021
Closest in time.
Seeing past words: Testing the cross-modal capabilities of pretrained v&l models on counting tasks
Letitia Parcalabescu, Albert Gatt, Anette Frank, and Iacer Calixto. 2021 · 2021
Closest in time.
On releasing annotator-level labels and information in datasets
Vinodkumar Prabhakaran, Aida Mostafazadeh Davani, and Mark Díaz. 2021 · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021 · 2021
Closest in time.
Are VQA systems RAD? Measuring robustness to augmented data with focused interventions
Daniel Rosenberg, Itai Gat, Amir Feder, and Roi Reichart. 2021 · 2021
Closest in time.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2031
Closest in time.