Fetching the paper…
Reading the bibliography…
'Actions' play a vital role in how humans interact with the world and enable them to achieve desired goals.
Programs with common sense
John McCarthy et al · 1960
Earlier work this paper cites.
Reasoning about actions and change: from single agent actions to multi-agent actions
Chitta Baral · 2010
Earlier work this paper cites.
The winograd schema challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern · 2012
Earlier work this paper cites.
Simulation as an engine of physical scene understanding
Peter W Battaglia, Jessica B Hamrick, and Joshua B Tenenbaum · 2013
Earlier work this paper cites.
Reporting bias and knowledge acquisition
Jonathan Gordon and Benjamin Van Durme · 2013
Earlier work this paper cites.
Commonsense reasoning and commonsense knowledge in artificial intelligence
Ernest Davis and Gary Marcus · 2015
Earlier work this paper cites.
Discovering states and transformations in image collections
Phillip Isola, Joseph Lim, and Edward Adelson · 2015
Earlier work this paper cites.
Towards ai-complete question answering: A set of prerequisite toy tasks, 2015
Jason Weston, Antoine Bordes, Sumit Chopra, Alexander M. Rush, Bart van Merriënboer, Armand Joulin, and Tomas Mikolov · 2015
Earlier work this paper cites.
Verb physics: Relative physical knowledge of actions and objects
Maxwell Forbes and Yejin Choi · 2017
Earlier work this paper cites.
Conceptnet 5.5: An open multilingual graph of general knowledge
Robyn Speer, Joshua Chin, and Catherine Havasi · 2017
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian D. Reid, Stephen Gould, and Anton van den Hengel · 2018
Earlier work this paper cites.
Learning interpretable spatial operations in a rich 3d blocks world
Yonatan Bisk, Kevin J. Shih, Yejin Choi, and Daniel Marcu · 2018
Earlier work this paper cites.
What action causes this? towards naive physical action-effect prediction
Qiaozi Gao, Shaohua Yang, Joyce Chai, and Lucy Vanderwende · 2018
Earlier work this paper cites.
Learning to describe differences between pairs of similar images
Harsh Jhamtani and Taylor Berg-Kirkpatrick · 2018
Earlier work this paper cites.
Answering visual what-if questions: From actions to predicted scene descriptions
Misha Wagner, Hector Basevi, Rakshith Shetty, Wenbin Li, Mateusz Malinowski, Mario Fritz, and Ales Leonardis · 2018
Earlier work this paper cites.
Recipeqa: A challenge dataset for multimodal comprehension of cooking recipes
Semih Yagcioglu, Aykut Erdem, Erkut Erdem, and Nazli Ikizler-Cinbis · 2018
Earlier work this paper cites.
Commonsense justification for action explanation
Shaohua Yang, Qiaozi Gao, Sari Saba-Sadiya, and Joyce Chai · 2018
Cited alongside, same era.
Visual choice of plausible alternatives: An evaluation of image-based commonsense causal reasoning
Jinyoung Yeo, Gyeongbok Lee, Gengyu Wang, Seungtaek Choi, Hyunsouk Cho, Reinald Kim Amplayo, and Seung-won Hwang · 2018
Cited alongside, same era.
Touchdown: Natural language navigation and spatial reasoning in visual street environments
Howard Chen, Alane Suhr, Dipendra Misra, Noah Snavely, and Yoav Artzi · 2019
Cited alongside, same era.
Everything happens for a reason: Discovering the purpose of actions in procedural text
Bhavana Dalvi, Niket Tandon, Antoine Bosselut, Wen-tau Yih, and Peter Clark · 2019
Cited alongside, same era.
Pre-learning environment representations for data-efficient neural instruction following
David Gaddy and Dan Klein · 2019
Cited alongside, same era.
Graph edit distance reward: Learning to edit scene graph
Lichang Chen, Guosheng Lin, Shijie Wang, and Qingyao Wu · 2020
Later among the works it cites.
Grounding physical concepts of objects and events through dynamic visual reasoning
Zhenfang Chen, Jiayuan Mao, Jiajun Wu, Kwan-Yee Kenneth Wong, Joshua B Tenenbaum, and Chuang Gan · 2020
Later among the works it cites.
Following instructions by imagining and reaching visual goals
John Kanu, Eadom Dessalene, Xiaomin Lin, Cornelia Fermuller, and Yiannis Aloimonos · 2020
Later among the works it cites.
Unifiedqa: Crossing format boundaries with a single qa system
Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi · 2020
Later among the works it cites.
Learning about objects by learning to interact with them
Martin Lohmann, Jordi Salvador, Aniruddha Kembhavi, and Roozbeh Mottaghi · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cooking with blocks : A recipe for visual reasoning on image-pairs
Tejas Gokhale, Shailaja Sampat, Zhiyuan Fang, Yezhou Yang, and Chitta Baral · 2019
Cited alongside, same era.
Vision-based navigation with language-based assistance via imitation learning with indirect intervention
Khanh Nguyen, Debadeepta Dey, Chris Brockett, and Bill Dolan · 2019
Cited alongside, same era.
Robust change captioning
Dong Huk Park, Trevor Darrell, and Anna Rohrbach · 2019
Cited alongside, same era.
Atomic: An atlas of machine commonsense for if-then reasoning
Maarten Sap, Ronan Le Bras, Emily Allaway, Chandra Bhagavatula, Nicholas Lourie, Hannah Rashkin, Brendan Roof, Noah A Smith, and Yejin Choi · 2019
Cited alongside, same era.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid · 2019
Cited alongside, same era.
Wiqa: A dataset for “what if…” reasoning over procedural text
Niket Tandon, Bhavana Dalvi, Keisuke Sakaguchi, Peter Clark, and Antoine Bosselut · 2019
Cited alongside, same era.
Composing text and image for image retrieval-an empirical odyssey
Nam Vo, Lu Jiang, Chen Sun, Kevin Murphy, Li-Jia Li, Li Fei-Fei, and James Hays · 2019
Cited alongside, same era.
Visualcomet: Reasoning about the dynamic context of a still image
Jae Sung Park, Chandra Bhagavatula, Roozbeh Mottaghi, Ali Farhadi, and Yejin Choi · 2020
Later among the works it cites.
Visuo-linguistic question answering (vlqa) challenge
Shailaja Keyur Sampat, Yezhou Yang, and Chitta Baral · 2020
Later among the works it cites.
Alfred: A benchmark for interpreting grounded instructions for everyday tasks
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox · 2020
Later among the works it cites.
Attention over learned object embeddings enables complex visual reasoning
David Ding, Felix Hill, Adam Santoro, Malcolm Reynolds, and Matt Botvinick · 2021
Later among the works it cites.
Transformation driven visual reasoning
Xin Hong, Yanyan Lan, Liang Pang, Jiafeng Guo, and Xueqi Cheng · 2021
Later among the works it cites.
Learning to compose visual relations
Nan Liu, Shuang Li, Yilun Du, Josh Tenenbaum, and Antonio Torralba · 2021
Later among the works it cites.
Clevr_hyp: A challenge dataset and baselines for visual question answering with hypothetical actions over images
Shailaja Keyur Sampat, Akshay Kumar, Yezhou Yang, and Chitta Baral · 2021
Later among the works it cites.
Tiered reasoning for intuitive physics: Toward verifiable commonsense language understanding
Shane Storks, Qiaozi Gao, Yichi Zhang, and Joyce Chai · 2021
Later among the works it cites.
Visual goal-step inference using wikihow
Yue Yang, Artemis Panagopoulou, Qing Lyu, Li Zhang, Mark Yatskar, and Chris Callison-Burch · 2021
Later among the works it cites.
The ai universe of “actions”: Agency, causality, commonsense and deception
Chitta Baral and Tran Cao Son · 2022
Closest in time.
A persistent spatial semantic representation for high-level natural language instruction execution
Valts Blukis, Chris Paxton, Dieter Fox, Animesh Garg, and Yoav Artzi · 2022
Closest in time.