Fetching the paper…
Reading the bibliography…
Videos often capture objects, their visible properties, their motion, and the interactions between different objects.
Monet: Unsupervised scene decomposition and representation
Christopher P. Burgess, Loïc Matthey, Nicholas Watters, Rishabh Kabra, Irina Higgins, Matthew Botvinick, and Alexander Lerchner. 2019 · 1901
Earlier work this paper cites.
A closer look at the robustness of vision-and-language pre-trained models
Linjie Li, Zhe Gan, and Jingjing Liu. 2020 · 2012
Earlier work this paper cites.
VQA: visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
A dataset for movie description
Anna Rohrbach, Marcus Rohrbach, Niket Tandon, and Bernt Schiele. 2015 · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015 · 2015
Earlier work this paper cites.
TGIF: A new dataset and benchmark on animated GIF description
Yuncheng Li, Yale Song, Liangliang Cao, Joel R. Tetreault, Larry Goldberg, Alejandro Jaimes, and Jiebo Luo. 2016 · 2016
Earlier work this paper cites.
Movieqa: Understanding stories in movies through question-answering
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. 2016 · 2016
Earlier work this paper cites.
MSR-VTT: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016 · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg. 2016 · 2016
Earlier work this paper cites.
Verb physics: Relative physical knowledge of actions and objects
Maxwell Forbes and Yejin Choi. 2017 · 2017
Earlier work this paper cites.
Mask R-CNN
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross B. Girshick. 2017 · 2017
Earlier work this paper cites.
TGIF-QA: toward spatio-temporal reasoning in visual question answering
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. 2017 · 2017
Earlier work this paper cites.
Inferring and executing programs for visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Judy Hoffman, Li Fei-Fei, C. Lawrence Zitnick, and Ross B. Girshick. 2017b · 2017
Earlier work this paper cites.
Compositional attention networks for machine reasoning
Drew A. Hudson and Christopher D. Manning. 2018 · 2018
Earlier work this paper cites.
Intphys: A framework and benchmark for visual intuitive physics reasoning
Ronan Riochet, Mario Ynocente Castro, Mathieu Bernard, Adam Lerer, Rob Fergus, Véronique Izard, and Emmanuel Dupoux. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Cater: A diagnostic dataset for compositional actions & temporal reasoning
Rohit Girdhar and Deva Ramanan. 2019 · 2019
Earlier work this paper cites.
Cooking with blocks: A recipe for visual reasoning on image-pairs
Tejas Gokhale, Shailaja Sampat, Zhiyuan Fang, Yezhou Yang, and Chitta Baral. 2019 · 2019
Cited alongside, same era.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning. 2019 · 2019
Cited alongside, same era.
The neuro-symbolic concept learner: Interpreting scenes, words, and sentences from natural supervision
Jiayuan Mao, Chuang Gan, Pushmeet Kohli, Joshua B. Tenenbaum, and Jiajun Wu. 2019 · 2019
Cited alongside, same era.
OK-VQA: A visual question answering benchmark requiring external knowledge
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. 2019 · 2019
Cited alongside, same era.
Sunny and dark outside?! improving answer consistency in VQA through entailed question generation
Arijit Ray, Karan Sikka, Ajay Divakaran, Stefan Lee, and Giedrius Burachas. 2019 · 2019
Cited alongside, same era.
Visualcomet: Reasoning about the dynamic context of a still image
Jae Sung Park, Chandra Bhagavatula, Roozbeh Mottaghi, Ali Farhadi, and Yejin Choi. 2020 · 2020
Later among the works it cites.
ESPRIT: Explaining solutions to physical reasoning tasks
Nazneen Fatema Rajani, Rui Zhang, Yi Chern Tan, Stephan Zheng, Jeremy Weiss, Aadit Vyas, Abhijit Gupta, Caiming Xiong, Richard Socher, and Dragomir Radev. 2020 · 2020
Later among the works it cites.
Visuo-linguistic question answering (vlqa) challenge
Shailaja Keyur Sampat, Yezhou Yang, and Chitta Baral. 2020 · 2020
Later among the works it cites.
Squinting at vqa models: Interrogating vqa models with sub-questions
Ramprasaath R. Selvaraju, Purva Tendulkar, Devi Parikh, Eric Horvitz, Marco Ribeiro, Besmira Nushi, and Ece Kamar. 2020 · 2020
Later among the works it cites.
CLEVRER: collision events for video representation and reasoning
Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, and Joshua B. Tenenbaum. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019 · 2019
Cited alongside, same era.
Composing text and image for image retrieval - an empirical odyssey
Nam Vo, Lu Jiang, Chen Sun, Kevin Murphy, Li-Jia Li, Li Fei-Fei, and James Hays. 2019 · 2019
Cited alongside, same era.
From recognition to cognition: Visual commonsense reasoning
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
Towards causal VQA: revealing and reducing spurious correlations by invariant and covariant semantic editing
Vedika Agarwal, Rakshith Shetty, and Mario Fritz. 2020 · 2020
Cited alongside, same era.
Cophy: Counterfactual learning of physical dynamics
Fabien Baradel, Natalia Neverova, Julien Mille, Greg Mori, and Christian Wolf. 2020 · 2020
Cited alongside, same era.
PIQA: reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan LeBras, Jianfeng Gao, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Procedure planning in instructional videos
Chien-Yi Chang, De-An Huang, Danfei Xu, Ehsan Adeli, Li Fei-Fei, and Juan Carlos Niebles. 2020 · 2020
Cited alongside, same era.
Attention over learned object embeddings enables complex visual reasoning
David Ding, Felix Hill, Adam Santoro, Malcolm Reynolds, and Matt Botvinick. 2021 · 2021
Later among the works it cites.
Threedworld: A platform for interactive multi-modal physical simulation
Chuang Gan, Jeremy Schwartz, Seth Alter, Damian Mrowca, Martin Schrimpf, James Traer, Julian De Freitas, Jonas Kubilius, Abhishek Bhandwaldar, Nick Haber, Megumi Sano, Kuno Kim, Elias Wang, Michael Lingelbach, Aidan Curtis, Kevin Feigelis, Daniel M. Bear, Dan Gutfreund, David Cox, Antonio Torralba, James J. DiCarlo, Joshua B. Tenenbaum, Josh H. McDermott, and Daniel L.K. Yamins. 2021 · 2021
Later among the works it cites.
Agqa: A benchmark for compositional spatio-temporal reasoning
Madeleine Grunde-McLaughlin, Ranjay Krishna, and Maneesh Agrawala. 2021 · 2021
Later among the works it cites.
Roses are red, violets are blue… but should vqa expect them to?
Corentin Kervadec, Grigory Antipov, Moez Baccouche, and Christian Wolf. 2021 · 2021
Later among the works it cites.
CLEVR_HYP: A challenge dataset and baselines for visual question answering with hypothetical actions over images
Shailaja Keyur Sampat, Akshay Kumar, Yezhou Yang, and Chitta Baral. 2021 · 2021
Later among the works it cites.
A case study of the shortcut effects in visual commonsense reasoning
K Ye and A Kovashka. 2021 · 2021
Later among the works it cites.
CRAFT: A benchmark for causal reasoning about forces and inTeractions
Tayfun Ates, M. Ateşoğlu, Çağatay Yiğit, Ilker Kesen, Mert Kobas, Erkut Erdem, Aykut Erdem, Tilbe Goksun, and Deniz Yuret. 2022 · 2022
Closest in time.
Comphy: Compositional physical reasoning of objects and events from videos
Zhenfang Chen, Kexin Yi, Yunzhu Li, Mingyu Ding, Antonio Torralba, Joshua B Tenenbaum, and Chuang Gan. 2022 · 2022
Closest in time.
Semantically distributed robust optimization for vision-and-language inference
Tejas Gokhale, Abhishek Chaudhary, Pratyay Banerjee, Chitta Baral, and Yezhou Yang. 2022 · 2022
Closest in time.
Reasoning about actions over visual and linguistic modalities: A survey
Shailaja Keyur Sampat, Maitreya Patel, Subhasish Das, Yezhou Yang, and Chitta Baral. 2022 · 2022
Closest in time.
Winoground: Probing vision and language models for visio-linguistic compositionality
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross. 2022 · 2022
Closest in time.