Fetching the paper…
Reading the bibliography…
Understanding human tasks through video observations is an essential capability of intelligent agents.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 1908
Earlier work this paper cites.
Situations, actions, and causal laws
John McCarthy · 1963
Earlier work this paper cites.
Strips: A new approach to the application of theorem proving to problem solving
Richard E Fikes and Nils J Nilsson · 1971
Earlier work this paper cites.
The frame problem in the situation calculus: A simple solution (sometimes) and a completeness result for goal regression
Raymond Reiter · 1991
Earlier work this paper cites.
Infants selectively encode the goal object of an actor’s reach
Amanda L Woodward · 1998
Earlier work this paper cites.
Pddl-the planning domain definition language
Drew McDermott, Malik Ghallab, Adele Howe, Craig Knoblock, Ashwin Ram, Manuela Veloso, Daniel Weld, and David Wilkins · 1998
Earlier work this paper cites.
The roles of vision and eye movements in the control of activities of daily living
Michael Land, Neil Mennie, and Jennifer Rusted · 1999
Earlier work this paper cites.
Infants parse dynamic action
Dare A Baldwin, Jodie A Baird, Megan M Saylor, and M Angela Clark · 2001
Earlier work this paper cites.
Rational imitation in preverbal infants
György Gergely, Harold Bekkering, and Ildikó Király · 2002
Earlier work this paper cites.
‘obsessed with goals’: Functions and mechanisms of teleological interpretation of actions in humans
Gergely Csibra and György Gergely · 2007
Earlier work this paper cites.
Action understanding as inverse planning
Chris L Baker, Rebecca Saxe, and Joshua B Tenenbaum · 2009
Earlier work this paper cites.
Detecting activities of daily living in first-person camera views
Hamed Pirsiavash and Deva Ramanan · 2012
Earlier work this paper cites.
Discovering localized attributes for fine-grained recognition
Kun Duan, Devi Parikh, David Crandall, and Kristen Grauman · 2012
Earlier work this paper cites.
Discovering important people and objects for egocentric video summarization
Yong Jae Lee, Joydeep Ghosh, and Kristen Grauman · 2012
Earlier work this paper cites.
Social interactions: A first-person perspective
Alircza Fathi, Jessica K Hodgins, and James M Rehg · 2012
Earlier work this paper cites.
Modeling actions through state changes
Alireza Fathi and James M Rehg · 2013
Earlier work this paper cites.
Story-driven summarization for egocentric video
Zheng Lu and Kristen Grauman · 2013
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
Hilde Kuehne, Ali Arslan, and Thomas Serre · 2014
Earlier work this paper cites.
Joint video and text parsing for understanding events and answering queries
Kewei Tu, Meng Meng, Mun Wai Lee, Tae Eun Choe, and Song-Chun Zhu · 2014
Earlier work this paper cites.
Discovering states and transformations in image collections
Phillip Isola, Joseph J Lim, and Edward H Adelson · 2015
Earlier work this paper cites.
Learning perceptual causality from video
Amy Fire and Song-Chun Zhu · 2015
Earlier work this paper cites.
Action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard L Lewis, and Satinder Singh · 2015
Earlier work this paper cites.
Lending a hand: Detecting hands and recognizing activities in complex egocentric interactions
Sven Bambach, Stefan Lee, David J Crandall, and Chen Yu · 2015
Earlier work this paper cites.
Social saliency prediction
Hyun Soo Park and Jianbo Shi · 2015
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Actions˜ transformations
Xiaolong Wang, Ali Farhadi, and Abhinav Gupta · 2016
Earlier work this paper cites.
Generating videos with scene dynamics
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Earlier work this paper cites.
Unsupervised learning from narrated instruction videos
Jean-Baptiste Alayrac, Piotr Bojanowski, Nishant Agrawal, Josef Sivic, Ivan Laptev, and Simon Lacoste-Julien · 2016
Earlier work this paper cites.
Functional object-oriented network for manipulation learning
David Paulius, Yongqiang Huang, Roger Milton, William D Buchanan, Jeanine Sam, and Yu Sun · 2016
Earlier work this paper cites.
Movieqa: Understanding stories in movies through question-answering
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler · 2016
Earlier work this paper cites.
You-do, i-learn: Egocentric unsupervised discovery of objects and their modes of interaction towards video-based guidance
Dima Damen, Teesid Leelasawassuk, and Walterio Mayol-Cuevas · 2016
Earlier work this paper cites.
Understanding hand-object manipulation with grasp types and object attributes
Minjie Cai, Kris M Kitani, and Yoichi Sato · 2016
Earlier work this paper cites.
Going deeper into first-person activity recognition
Minghuang Ma, Haoqi Fan, and Kris M Kitani · 2016
Earlier work this paper cites.
Joint discovery of object states and manipulation actions
Jean-Baptiste Alayrac, Ivan Laptev, Josef Sivic, and Simon Lacoste-Julien · 2017
Earlier work this paper cites.
Jointly recognizing object fluents and tasks in egocentric videos
Yang Liu, Ping Wei, and Song-Chun Zhu · 2017
Cited alongside, same era.
Marioqa: Answering questions by watching gameplay videos
Jonghwan Mun, Paul Hongsuck Seo, Ilchae Jung, and Bohyung Han · 2017
Cited alongside, same era.
Deepstory: Video story QA by deep embedded memory networks
Kyung-Min Kim, Min-Oh Heo, Seong-Ho Choi, and Byoung-Tak Zhang · 2017
Cited alongside, same era.
Video question answering via gradually refined attention over appearance and motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang · 2017
Cited alongside, same era.
Leveraging video descriptions to learn video question answering
Kuo-Hao Zeng, Tseng-Hung Chen, Ching-Yao Chuang, Yuan-Hong Liao, Juan Carlos Niebles, and Min Sun · 2017
Cited alongside, same era.
TGIF-QA: toward spatio-temporal reasoning in visual question answering
TVQA+: spatio-temporal grounding for video question answering
Jie Lei, Licheng Yu, Tamara L. Berg, and Mohit Bansal · 2020
Later among the works it cites.
HERO: hierarchical encoder for video+language omni-representation pre-training
Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu · 2020
Later among the works it cites.
Ego-topo: Environment affordances from egocentric video
Tushar Nagarajan, Yanghao Li, Christoph Feichtenhofer, and Kristen Grauman · 2020
Later among the works it cites.
Rolling-unrolling lstms for action anticipation from first-person video
Antonino Furnari and Giovanni Maria Farinella · 2020
Later among the works it cites.
A generalized earley parser for human activity parsing and prediction
Siyuan Qi, Baoxiong Jia, Siyuan Huang, Ping Wei, and Song-Chun Zhu · 2020
Later among the works it cites.
You2me: Inferring body pose in egocentric video via first and second person interactions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim · 2017
Cited alongside, same era.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Cited alongside, same era.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Cited alongside, same era.
Charades-ego: A large-scale dataset of paired third and first person videos
Gunnar A Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari · 2018
Cited alongside, same era.
In the eye of beholder: Joint learning of gaze and actions in first person video
Yin Li, Miao Liu, and James M Rehg · 2018
Cited alongside, same era.
Attributes as operators: factorizing unseen attribute-object compositions
Tushar Nagarajan and Kristen Grauman · 2018
Cited alongside, same era.
Recurrent world models facilitate policy evolution
David Ha and Jürgen Schmidhuber · 2018
Cited alongside, same era.
Evonne Ng, Donglai Xiang, Hanbyul Joo, and Kristen Grauman · 2020
Later among the works it cites.
Egocom: A multi-person multi-modal egocentric communications dataset
Curtis Northcutt, Shengxin Zha, Steven Lovegrove, and Richard Newcombe · 2020
Later among the works it cites.
Bongard-logo: A new benchmark for human-level concept learning and reasoning
Weili Nie, Zhiding Yu, Lei Mao, Ankit B Patel, Yuke Zhu, and Anima Anandkumar · 2020
Later among the works it cites.
Visualcomet: Reasoning about the dynamic context of a still image
Jae Sung Park, Chandra Bhagavatula, Roozbeh Mottaghi, Ali Farhadi, and Yejin Choi · 2020
Later among the works it cites.
Action genome: Actions as compositions of spatio-temporal scene graphs
Jingwei Ji, Ranjay Krishna, Li Fei-Fei, and Juan Carlos Niebles · 2020
Later among the works it cites.
Reasoning with heterogeneous graph alignment for video question answering
Pin Jiang and Yahong Han · 2020
Later among the works it cites.
Hierarchical conditional relation networks for video question answering
Thao Minh Le, Vuong Le, Svetha Venkatesh, and Truyen Tran · 2020
Later among the works it cites.
In defense of grid features for visual question answering
Huaizu Jiang, Ishan Misra, Marcus Rohrbach, Erik Learned-Miller, and Xinlei Chen · 2020
Later among the works it cites.
Learning triadic belief dynamics in nonverbal communication from videos
Lifeng Fan, Shuwen Qiu, Zilong Zheng, Tao Gao, Song-Chun Zhu, and Yixin Zhu · 2021
Later among the works it cites.
Piglet: Language grounding through neuro-symbolic interaction in a 3d world
Rowan Zellers, Ari Holtzman, Matthew Peters, Roozbeh Mottaghi, Aniruddha Kembhavi, Ali Farhadi, and Yejin Choi · 2021
Later among the works it cites.
Learning temporal dynamics from cycles in narrated video
Dave Epstein, Jiajun Wu, Cordelia Schmid, and Chen Sun · 2021
Later among the works it cites.
Env-qa: A video question answering benchmark for comprehensive understanding of dynamic environments
Difei Gao, Ruiping Wang, Ziyi Bai, and Xilin Chen · 2021
Later among the works it cites.
Just ask: Learning to answer questions from millions of narrated videos
Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid · 2021
Later among the works it cites.
AGQA: A benchmark for compositional spatio-temporal reasoning
Madeleine Grunde-McLaughlin, Ranjay Krishna, and Maneesh Agrawala · 2021
Later among the works it cites.
Next-qa: Next phase of question-answering to explaining temporal actions
Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua · 2021
Later among the works it cites.
Star: A benchmark for situated reasoning in real-world videos
Bo Wu, Shoubin Yu, Zhenfang Chen, Joshua B Tenenbaum, and Chuang Gan · 2021
Later among the works it cites.
Anticipative video transformer
Rohit Girdhar and Kristen Grauman · 2021
Later among the works it cites.
Home action genome: Cooperative compositional action understanding
Nishant Rai, Haofeng Chen, Jingwei Ji, Rishi Desai, Kazuki Kozuka, Shun Ishizaka, Ehsan Adeli, and Juan Carlos Niebles · 2021
Later among the works it cites.
Less is more: Clipbert for video-and-language learning via sparse sampling
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu · 2021
Later among the works it cites.
R3M: A universal visual representation for robot manipulation
Suraj Nair, Aravind Rajeswaran, Vikash Kumar, Chelsea Finn, and Abhinav Gupta · 2022
Closest in time.
Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al · 2022
Closest in time.
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al · 2022
Closest in time.
Teach: Task-driven embodied agents that chat
Aishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava, Patrick Lange, Anjali Narayan-Chen, Spandana Gella, Robinson Piramuthu, Gokhan Tur, and Dilek Hakkani-Tur · 2022
Closest in time.
Plate: Visually-grounded planning with transformers in procedural tasks
Jiankai Sun, De-An Huang, Bo Lu, Yun-Hui Liu, Bolei Zhou, and Animesh Garg · 2022
Closest in time.
Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition
Chao-Yuan Wu, Yanghao Li, Karttikeya Mangalam, Haoqi Fan, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer · 2022
Closest in time.
Omnivore: A Single Model for Many Visual Modalities
Rohit Girdhar, Mannat Singh, Nikhila Ravi, Laurens van der Maaten, Armand Joulin, and Ishan Misra · 2022
Closest in time.
Learning algebraic representation for systematic generalization in abstract reasoning
Chi Zhang, Sirui Xie, Baoxiong Jia, Ying Nian Wu, Song-Chun Zhu, and Yixin Zhu · 2022
Closest in time.
Learning to recognize procedural activities with distant supervision
Xudong Lin, Fabio Petroni, Gedas Bertasius, Marcus Rohrbach, Shih-Fu Chang, and Lorenzo Torresani · 2022
Closest in time.