Fetching the paper…
Reading the bibliography…
Visual commonsense plays a vital role in understanding and reasoning about the visual world.
The reviewing of object files: Object-specific integration of information
Daniel Kahneman, Anne Treisman, and Brian J Gibbs. 1992 · 1992
Earlier work this paper cites.
A new quantitative quality measure for machine translation systems
Keh-Yih Su, Ming-Wen Wu, and Jing-Shin Chang. 1992 · 1992
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Visualsem: a high-quality knowledge graph for vision and language
Houda Alberts, Teresa Huang, Yash Deshpande, Yibo Liu, Kyunghyun Cho, Clara Vania, and Iacer Calixto. 2020 · 2008
Earlier work this paper cites.
Neil: Extracting visual knowledge from web data
Xinlei Chen, Abhinav Shrivastava, and Abhinav Gupta. 2013 · 2013
Earlier work this paper cites.
Mining semantic affordances of visual object categories
Yu-Wei Chao, Zhan Wang, Rada Mihalcea, and Jia Deng. 2015 · 2015
Earlier work this paper cites.
Learning common sense through visual abstraction
Ramakrishna Vedantam, Xiao Lin, Tanmay Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Stating the obvious: Extracting visual common sense knowledge
Mark Yatskar, Vicente Ordonez, and Ali Farhadi. 2016 · 2016
Earlier work this paper cites.
Imgpedia: a linked dataset with content-based analysis of wikimedia images
Sebastián Ferrada, Benjamin Bustos, and Aidan Hogan. 2017 · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017 · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei. 2017 · 2017
Earlier work this paper cites.
Answering visual-relational queries in web-extracted knowledge graphs
Daniel Oñoro-Rubio, Mathias Niepert, Alberto García-Durán, Roberto González, and Roberto J López-Sastre. 2017 · 2017
Earlier work this paper cites.
Conceptnet 5.5: An open multilingual graph of general knowledge
Robyn Speer, Joshua Chin, and Catherine Havasi. 2017 · 2017
Earlier work this paper cites.
Visual relationship detection with internal and external linguistic knowledge distillation
Ruichi Yu, Ang Li, Vlad I. Morariu, and Larry S. Davis. 2017 · 2017
Earlier work this paper cites.
Acquiring common sense spatial knowledge through implicit spatial templates
Guillem Collell, Luc Van Gool, and Marie-Francine Moens. 2018 · 2018
Earlier work this paper cites.
AllenNLP: A deep semantic natural language processing platform
Matt Gardner, Joel Grus, Mark Neumann, Oyvind Tafjord, Pradeep Dasigi, Nelson F. Liu, Matthew Peters, Michael Schmitz, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Automatic extraction of commonsense LocatedNear knowledge
Frank F. Xu, Bill Yuchen Lin, and Kenny Zhu. 2018 · 2018
Cited alongside, same era.
Mmkg: multi-modal knowledge graphs
Ye Liu, Hui Li, Alberto Garcia-Duran, Mathias Niepert, Daniel Onoro-Rubio, and David S Rosenblum. 2019 · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Cited alongside, same era.
OK-VQA: A visual question answering benchmark requiring external knowledge
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. 2019 · 2019
Cited alongside, same era.
Atomic: An atlas of machine commonsense for if-then reasoning
Maarten Sap, Ronan Le Bras, Emily Allaway, Chandra Bhagavatula, Nicholas Lourie, Hannah Rashkin, Brendan Roof, Noah A Smith, and Yejin Choi. 2019 · 2019
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Later among the works it cites.
LAION-5B: an open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev. 2022 · 2022
Later among the works it cites.
OFA: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. 2022 · 2022
Later among the works it cites.
Visual commonsense in pretrained unimodal and multimodal models
Chenyu Zhang, Benjamin Van Durme, Zhuowan Li, and Elias Stengel-Eskin. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
From recognition to cognition: Visual commonsense reasoning
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
spaCy: Industrial-strength Natural Language Processing in Python
Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd. 2020 · 2020
Cited alongside, same era.
Visualcomet: Reasoning about the dynamic context of a still image
Jae Sung Park, Chandra Bhagavatula, Roozbeh Mottaghi, Ali Farhadi, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Richpedia: a large-scale, comprehensive multi-modal knowledge graph
Meng Wang, Haofen Wang, Guilin Qi, and Qiushuo Zheng. 2020 · 2020
Cited alongside, same era.
Transomcs: From linguistic graphs to commonsense knowledge
Hongming Zhang, Daniel Khashabi, Yangqiu Song, and Dan Roth. 2020a · 2020
Cited alongside, same era.
ASER: A large-scale eventuality knowledge graph
Hongming Zhang, Xin Liu, Haojie Pan, Yangqiu Song, and Cane Wing-Ki Leung. 2020b · 2020
Cited alongside, same era.
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven C. H. Hoi. 2023 · 2023
Later among the works it cites.
Mme: A comprehensive evaluation benchmark for multimodal large language models
Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Zhenyu Qiu, Wei Lin, Jinrui Yang, Xiawu Zheng, et al. 2023 · 2023
Later among the works it cites.
Datacomp: In search of the next generation of multimodal datasets
Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, Eyal Orgad, Rahim Entezari, Giannis Daras, Sarah M. Pratt, Vivek Ramanujan, Yonatan Bitton, Kalyani Marathe, Stephen Mussmann, Richard Vencu, Mehdi Cherti, Ranjay Krishna, Pang Wei Koh, Olga Saukh, Alexander J. Ratner, Shuran Song, Hannaneh Hajishirzi, Ali Farhadi, Romain Beaumont, Sewoong Oh, Alex Dimakis, Jenia Jitsev, Yair Carmon, Vaishaal Shankar, and Ludwig Schmidt. 2023 · 2023
Later among the works it cites.
Can language models understand physical concepts?
Lei Li, Jingjing Xu, Qingxiu Dong, Ce Zheng, Xu Sun, Lingpeng Kong, and Qi Liu. 2023 · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
Dense-atomic: Towards densely-connected ATOMIC with high knowledge coverage and massive multi-hop paths
Xiangqing Shen, Siwei Wu, and Rui Xia. 2023 · 2023
Later among the works it cites.
VIPHY: Probing “visible” physical commonsense knowledge
Shikhar Singh, Ehsan Qasemi, and Muhao Chen. 2023 · 2023
Later among the works it cites.
Intrinsic physical concepts discovery with object-centric predictive models
Qu Tang, Xiangyu Zhu, Zhen Lei, and Zhaoxiang Zhang. 2023 · 2023
Later among the works it cites.
Imagenetvc: Zero- and few-shot visual commonsense evaluation on 1000 imagenet categories
Heming Xia, Qingxiu Dong, Lei Li, Jingjing Xu, Tianyu Liu, Ziwei Qin, and Zhifang Sui. 2023 · 2023
Later among the works it cites.
MultiInstruct: Improving multi-modal zero-shot learning via instruction tuning
Zhiyang Xu, Ying Shen, and Lifu Huang. 2023 · 2023
Later among the works it cites.
ViCor: Bridging visual understanding and commonsense reasoning with large language models
Kaiwen Zhou, Kwonjoon Lee, Teruhisa Misu, and Xin Wang. 2024 · 2024
Closest in time.
PIGLeT: Language grounding through neuro-symbolic interaction in a 3D world
Rowan Zellers, Ari Holtzman, Matthew Peters, Roozbeh Mottaghi, Aniruddha Kembhavi, Ali Farhadi, and Yejin Choi. 2021 · 2050
Closest in time.