Fetching the paper…
Reading the bibliography…
In this paper, we propose a Disentangled Counterfactual Learning~(DCL) approach for physical audiovisual commonsense reasoning.
Qualitative reasoning about physical systems: an introduction
Daniel G Bobrow · 1984
Earlier work this paper cites.
Qualitative process theory
Kenneth D Forbus · 1984
Earlier work this paper cites.
Framewise phoneme classification with bidirectional lstm and other neural network architectures
Alex Graves and Jürgen Schmidhuber · 2005
Earlier work this paper cites.
Shake, rattle, and… one or two objects? young infants’ use of auditory information to individuate objects
Teresa Wilcox, Rebecca Woods, Lisa Tuggy, and Roman Napoli · 2006
Earlier work this paper cites.
Making sense of virtual environments: action representation, grounding and common sense
Jean-Luc Lugrin and Marc Cavazza · 2007
Earlier work this paper cites.
Commonsense reasoning about the physical world
Joan Bliss · 2008
Earlier work this paper cites.
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Deep neural networks: a new framework for modeling biological vision and brain information processing
Nikolaus Kriegeskorte · 2015
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Five-month-old infants have general knowledge of how nonsolid substances behave and interact
Susan J Hespos, Alissa L Ferry, Erin M Anderson, Emily N Hollenbeck, and Lance J Rips · 2016
Earlier work this paper cites.
Causal inference in statistics: A primer
Madelyn Glymour, Judea Pearl, and Nicholas P Jewell · 2016
Earlier work this paper cites.
An unsupervised neural attention model for aspect extraction
Ruidan He, Wee Sun Lee, Hwee Tou Ng, and Daniel Dahlmeier · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
Diverse image-to-image translation via disentangled representations
Hsin-Ying Lee, Hung-Yu Tseng, Jia-Bin Huang, Maneesh Singh, and Ming-Hsuan Yang · 2018
Earlier work this paper cites.
Visual object networks: Image generation with disentangled 3d representations
Jun-Yan Zhu, Zhoutong Zhang, Chengkai Zhang, Jiajun Wu, Antonio Torralba, Josh Tenenbaum, and Bill Freeman · 2018
Earlier work this paper cites.
Grounding visual explanations
Lisa Anne Hendricks, Ronghang Hu, Trevor Darrell, and Zeynep Akata · 2018
Cited alongside, same era.
stagnet: An attentive semantic rnn for group activity recognition
Mengshi Qi, Jie Qin, Annan Li, Yunhong Wang, Jiebo Luo, and Luc Van Gool · 2018
Cited alongside, same era.
Disentangled sequential autoencoder
Li Yingzhen and Stephan Mandt · 2018
Cited alongside, same era.
The book of why: the new science of cause and effect
Judea Pearl and Dana Mackenzie · 2018
Cited alongside, same era.
Do neural language representations learn physical commonsense?
Maxwell Forbes, Ari Holtzman, and Yejin Choi · 2019
Cited alongside, same era.
Generating sentences from disentangled syntactic and semantic spaces
Yu Bao, Hao Zhou, Shujian Huang, Lei Li, Lili Mou, Olga Vechtomova, Xin-yu Dai, and Jiajun Chen · 2019
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2020
Later among the works it cites.
Audio-visual floorplan reconstruction
Senthil Purushwalkam, Sebastia Vicenc Amengual Gari, Vamsi Krishna Ithapu, Carl Schissler, Philip Robinson, Abhinav Gupta, and Kristen Grauman · 2021
Later among the works it cites.
Semantics-aware spatial-temporal binaries for cross-modal video retrieval
Mengshi Qi, Jie Qin, Yi Yang, Yunhong Wang, and Jiebo Luo · 2021
Later among the works it cites.
Pano-avqa: Grounded audio-visual question answering on 360deg videos
Heeseung Yun, Youngjae Yu, Wonsuk Yang, Kangil Lee, and Gunhee Kim · 2021
Later among the works it cites.
Contrastively disentangled sequential variational autoencoder
Junwen Bai, Weiran Wang, and Carla P Gomes · 2021
Later among the works it cites.
Indigo: Gnn-based inductive knowledge graph completion using pair-wise encoding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Counterfactual visual explanations
Yash Goyal, Ziyan Wu, Jan Ernst, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Cited alongside, same era.
A corpus for reasoning about natural language grounded in photographs
Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi · 2019
Cited alongside, same era.
Robot action planning by commonsense knowledge in human-robot collaborative tasks
Christopher J Conti, Aparna S Varde, and Weitian Wang · 2020
Cited alongside, same era.
Look, listen, and act: Towards audio-visual embodied navigation
Chuang Gan, Yiwei Zhang, Jiajun Wu, Boqing Gong, and Joshua B Tenenbaum · 2020
Cited alongside, same era.
Piqa: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al · 2020
Cited alongside, same era.
Scout: Self-aware discriminant counterfactual explanations
Pei Wang and Nuno Vasconcelos · 2020
Cited alongside, same era.
Shuwen Liu, Bernardo Grau, Ian Horrocks, and Egor Kostylev · 2021
Later among the works it cites.
Deep learning-based late fusion of multimodal information for emotion classification of music video
Yagya Raj Pandeya and Joonwhoan Lee · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Pacs: A dataset for physical audiovisual commonsense reasoning
Samuel Yu, Peter Wu, Paul Pu Liang, Ruslan Salakhutdinov, and Louis-Philippe Morency · 2022
Later among the works it cites.
Finding fallen objects via asynchronous audio-visual integration
Chuang Gan, Yi Gu, Siyuan Zhou, Jeremy Schwartz, Seth Alter, James Traer, Dan Gutfreund, Joshua B Tenenbaum, Josh H McDermott, and Antonio Torralba · 2022
Later among the works it cites.
Mukea: Multimodal knowledge extraction and accumulation for knowledge-based visual question answering
Yang Ding, Jing Yu, Bang Liu, Yue Hu, Mingxin Cui, and Qi Wu · 2022
Later among the works it cites.
Direct and indirect effects
Judea Pearl · 2022
Later among the works it cites.
Audioclip: Extending clip to image, text and audio
Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas Dengel · 2022
Later among the works it cites.
Merlot reserve: Neural script knowledge through vision and language and sound
Rowan Zellers, Jiasen Lu, Ximing Lu, Youngjae Yu, Yanpeng Zhao, Mohammadreza Salehi, Aditya Kusupati, Jack Hessel, Ali Farhadi, and Yejin Choi · 2022
Later among the works it cites.
Progressive spatio-temporal perception for audio-visual question answering
Guangyao Li, Wenxuan Hou, and Di Hu · 2023
Closest in time.