Fetching the paper…
Reading the bibliography…
While vision-and-language models perform well on tasks such as visual question answering, they struggle when it comes to basic human commonsense reasoning skills.
Do neural language representations learn physical commonsense?
Maxwell Forbes, Ari Holtzman, and Yejin Choi · 1908
Earlier work this paper cites.
Gender bias in coreference resolution
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme · 2002
Earlier work this paper cites.
A review of winograd schema challenge datasets and approaches
Vid Kocijan, Thomas Lukasiewicz, Ernest Davis, Gary Marcus, and Leora Morgenstern · 2004
Earlier work this paper cites.
Labeling images with a computer game
Luis Von Ahn and Laura Dabbish · 2004
Earlier work this paper cites.
SWOW-8500: Word association task for intrinsic evaluation of word embeddings
Avijit Thawani, Biplav Srivastava, and Anil Singh · 2006
Earlier work this paper cites.
Wikirelate! computing semantic relatedness using wikipedia
Michael Strube and Simone Paolo Ponzetto · 2006
Earlier work this paper cites.
Evaluating WordNet-based measures of lexical semantic relatedness
Alexander Budanitsky and Graeme Hirst · 2006
Earlier work this paper cites.
Computing semantic relatedness using wikipedia-based explicit semantic analysis
Evgeniy Gabrilovich, Shaul Markovitch, et al · 2007
Earlier work this paper cites.
Semantic relatedness using salient semantic analysis
Samer Hassan and Rada Mihalcea · 2011
Earlier work this paper cites.
The winograd schema challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern · 2012
Earlier work this paper cites.
Quizz: targeted crowdsourcing with a billion (potential) users
Panagiotis G. Ipeirotis and Evgeniy Gabrilovich · 2014
Earlier work this paper cites.
Concreteness ratings for 40 thousand generally known english word lemmas
Marc Brysbaert, Amy Beth Warriner, and Victor Kuperman · 2014
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Learning common sense through visual abstraction
Ramakrishna Vedantam, Xiao Lin, Tanmay Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Beat the machine: Challenging humans to find a predictive model’s “unknown unknowns”
Joshua Attenberg, Panos Ipeirotis, and Foster Provost · 2015
Earlier work this paper cites.
Explicit retrieval of visual and non-visual properties of concrete entities
Antonietta Gabriella Liuzzi, Patrick Dupont, Ronald Peeters, Simon De Deyne, Gerrit Storms, and Rik Vandenberghe · 2017
Earlier work this paper cites.
Large-scale network representations of semantics in the mental lexicon
Simon De Deyne, Yoed N Kenett, David Anaki, Miriam Faust, and Daniel Navarro · 2017
Earlier work this paper cites.
Game on! experiential learning with tabletop games
L Hays and M Hayse · 2017
Earlier work this paper cites.
Evaluating visual conversational agents via cooperative human-ai games
Prithvijit Chattopadhyay, Deshraj Yadav, Viraj Prabhu, Arjun Chandrasekaran, Abhishek Das, Stefan Lee, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
Visual and affective grounding in language and mind
Simon De Deyne, Danielle Navarro, Guillem Collell, and Andrew Perfors · 2018
Earlier work this paper cites.
Comparing models of associative meaning: An empirical investigation of reference in simple language games
Judy Hanwen Shen, Matthias Hofer, Bjarke Felbo, and Roger Levy · 2018
Earlier work this paper cites.
SWAG: A large-scale adversarial dataset for grounded commonsense inference
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi · 2018
Cited alongside, same era.
New perspectives on the aging lexicon
Dirk U Wulff, Simon De Deyne, Michael N Jones, Rui Mata, Aging Lexicon Consortium, et al · 2019
Cited alongside, same era.
From recognition to cognition: Visual commonsense reasoning
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Cited alongside, same era.
The “small world of words” english word association norms for over 12,000 cue words
Simon De Deyne, Danielle J Navarro, Amy Perfors, Marc Brysbaert, and Gert Storms · 2019
Cited alongside, same era.
Character region awareness for text detection
Youngmin Baek, Bado Lee, Dongyoon Han, Sangdoo Yun, and Hwalsuk Lee · 2019
Cited alongside, same era.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Nils Reimers and Iryna Gurevych · 2019
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Later among the works it cites.
How much can clip benefit vision-and-language tasks?
Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal, Anna Rohrbach, Kai-Wei Chang, Zhewei Yao, and Kurt Keutzer · 2021
Later among the works it cites.
Vilt: Vision-and-language transformer without convolution or region supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim · 2021
Later among the works it cites.
Multi-grained vision language pre-training: Aligning texts with visual concepts
Yan Zeng, Xinsong Zhang, and Hang Li · 2021
Later among the works it cites.
Wino-X: Multilingual Winograd schemas for commonsense reasoning and coreference resolution
Denis Emelin and Rico Sennrich · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Cited alongside, same era.
Cooperation and codenames: Understanding natural language processing via codenames
Andrew Kim, Maxim Ruzmaykin, Aaron Truong, and Adam Summerville · 2019
Cited alongside, same era.
ATOMIC: an atlas of machine commonsense for if-then reasoning
Maarten Sap, Ronan Le Bras, Emily Allaway, Chandra Bhagavatula, Nicholas Lourie, Hannah Rashkin, Brendan Roof, Noah A. Smith, and Yejin Choi · 2019
Cited alongside, same era.
Mpnet: Masked and permuted pre-training for language understanding
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu · 2020
Cited alongside, same era.
PIQA: reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan LeBras, Jianfeng Gao, and Yejin Choi · 2020
Cited alongside, same era.
Video2Commonsense: Generating commonsense descriptions to enrich video captioning
Zhiyuan Fang, Tejas Gokhale, Pratyay Banerjee, Chitta Baral, and Yezhou Yang · 2020
Cited alongside, same era.
Later among the works it cites.
Abstraction and analogy-making in artificial intelligence
Melanie Mitchell · 2021
Later among the works it cites.
ExplaGraphs: An explanation graph generation task for structured commonsense reasoning
Swarnadeep Saha, Prateek Yadav, Lisa Bauer, and Mohit Bansal · 2021
Later among the works it cites.
Adversarial vqa: A new benchmark for evaluating the robustness of vqa models
Linjie Li, Jie Lei, Zhe Gan, and Jingjing Liu · 2021
Later among the works it cites.
Human-adversarial visual question answering
Sasha Sheng, Amanpreet Singh, Vedanuj Goswami, Jose Magana, Tristan Thrush, Wojciech Galuba, Devi Parikh, and Douwe Kiela · 2021
Later among the works it cites.
Fool me twice: Entailment from Wikipedia gamification
Julian Eisenschlos, Bhuwan Dhingra, Jannis Bulian, Benjamin Börschinger, and Jordan Boyd-Graber · 2021
Later among the works it cites.
Cultural and Academic Experiences of Black Male Graduates from a Historically Black College and University
Charea Lacherie Bustamante · 2021
Later among the works it cites.
Dynabench: Rethinking benchmarking in NLP
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Later among the works it cites.
Winoground: Probing vision and language models for visio-linguistic compositionality
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross · 2022
Closest in time.
Jack Hessel, Ana Marasović, Jena D Hwang, Lillian Lee, Jeff Da, Rowan Zellers, Robert Mankoff, and Yejin Choi · 2022
Closest in time.
Commonsenseqa 2.0: Exposing the limits of ai through gamification
Alon Talmor, Ori Yoran, Ronan Le Bras, Chandra Bhagavatula, Yoav Goldberg, Yejin Choi, and Jonathan Berant · 2022
Closest in time.
Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang · 2022
Closest in time.
The curious case of commonsense intelligence
Yejin Choi · 2022
Closest in time.
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie · 2022
Closest in time.