Fetching the paper…
Reading the bibliography…
Object State Changes (OSCs) are pivotal for video understanding.
Toward open set recognition
Walter J Scheirer, Anderson de Rezende Rocha, Archana Sapkota, and Terrance E Boult · 2012
Earlier work this paper cites.
Modeling actions through state changes
Alireza Fathi and James M Rehg · 2013
Earlier work this paper cites.
Towards open world recognition
Abhijit Bendale and Terrance Boult · 2015
Earlier work this paper cites.
Discovering states and transformations in image collections
Phillip Isola, Joseph J. Lim, and Edward H. Adelson · 2015
Earlier work this paper cites.
Joint discovery of object states and manipulation actions
Jean-Baptiste Alayrac, Ivan Laptev, Josef Sivic, and Simon Lacoste-Julien · 2017
Earlier work this paper cites.
Detecting the moment of completion: Temporal models for localising action completion
Farnoosh Heidarivincheh, Majid Mirmehdi, and Dima Damen · 2017
Earlier work this paper cites.
Jointly recognizing object fluents and tasks in egocentric videos
Yang Liu, Ping Wei, and Song-Chun Zhu · 2017
Earlier work this paper cites.
From red wine to red tomato: Composition with context
Ishan Misra, Abhinav Gupta, and Martial Hebert · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Action completion: A temporal model for moment detection
Farnoosh Heidarivincheh, Majid Mirmehdi, and Dima Damen · 2018
Earlier work this paper cites.
Attributes as operators: factorizing unseen attribute-object compositions
Tushar Nagarajan and Kristen Grauman · 2018
Earlier work this paper cites.
Deep object pose estimation for semantic robotic grasping of household objects
Jonathan Tremblay, Thang To, Balakumar Sundaralingam, Yu Xiang, Dieter Fox, and Stan Birchfield · 2018
Earlier work this paper cites.
Recognizing manipulation actions from state-transformations
Nachwa Aboubakr, James L Crowley, and Rémi Ronfard · 2019
Earlier work this paper cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Earlier work this paper cites.
Task-driven modular networks for zero-shot compositional learning
Senthil Purushwalkam, Maximilian Nickel, Abhinav Gupta, and Marc’Aurelio Ranzato · 2019
Earlier work this paper cites.
Procedure planning in instructional videos
Chien-Yi Chang, De-An Huang, Danfei Xu, Ehsan Adeli, Li Fei-Fei, and Juan Carlos Niebles · 2020
Earlier work this paper cites.
Deep learning in video multi-object tracking: A survey
Gioele Ciaparrone, Francisco Luque Sánchez, Siham Tabik, Luigi Troiano, Roberto Tagliaferri, and Francisco Herrera · 2020
Earlier work this paper cites.
Self-supervised 6d object pose estimation for robot manipulation
Xinke Deng, Yu Xiang, Arsalan Mousavian, Clemens Eppner, Timothy Bretl, and Dieter Fox · 2020
Earlier work this paper cites.
The overlooked elephant of object detection: Open set
Akshay Dhamija, Manuel Gunther, Jonathan Ventura, and Terrance Boult · 2020
Earlier work this paper cites.
Oops! predicting unintentional action in video
Dave Epstein, Boyuan Chen, and Carl Vondrick · 2020
Earlier work this paper cites.
Something-else: Compositional action recognition with spatial-temporal interaction networks
Joanna Materzynska, Tete Xiao, Roei Herzig, Huijuan Xu, Xiaolong Wang, and Trevor Darrell · 2020
Cited alongside, same era.
Understanding human hands in contact at internet scale
Dandan Shan, Jiaqi Geng, Michelle Shu, and David F Fouhey · 2020
Cited alongside, same era.
Video object segmentation and tracking: A survey
Rui Yao, Guosheng Lin, Shixiong Xia, Jiaqi Zhao, and Yong Zhou · 2020
Cited alongside, same era.
Procedure planning in instructional videos via contextual modeling and model-based policy learning
Jing Bi, Jiebo Luo, and Chenliang Xu · 2021
Cited alongside, same era.
Learning goals from failure
Dave Epstein and Carl Vondrick · 2021
Cited alongside, same era.
Learning temporal dynamics from cycles in narrated video
Dave Epstein, Jiajun Wu, Cordelia Schmid, and Chen Sun · 2021
Prompting visual-language models for efficient video understanding
Chen Ju, Tengda Han, Kunhao Zheng, Ya Zhang, and Weidi Xie · 2022
Later among the works it cites.
Egocentric video-language pretraining
Kevin Qinghong Lin, Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Z XU, Difei Gao, Rong-Cheng Tu, Wenzhe Zhao, Weijie Kong, et al · 2022
Later among the works it cites.
Image segmentation using text and image prompts
Timo Lüddecke and Alexander Ecker · 2022
Later among the works it cites.
Recognizing actions using object states
Nirat Saini, Bo He, Gaurav Shrivastava, Sai Saketh Rambhatla, and Abhinav Shrivastava · 2022
Later among the works it cites.
Internvideo: General video foundation models via generative and discriminative learning
Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang, Jilan Xu, Yi Liu, Zun Wang, et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Temporal roi align for video object recognition
Tao Gong, Kai Chen, Xinjiang Wang, Qi Chu, Feng Zhu, Dahua Lin, Nenghai Yu, and Huamin Feng · 2021
Cited alongside, same era.
Open-vocabulary object detection via vision and language knowledge distillation
Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo, and Yin Cui · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig · 2021
Cited alongside, same era.
New generation deep learning for video object detection: A survey
Licheng Jiao, Ruohan Zhang, Fang Liu, Shuyuan Yang, Biao Hou, Lingling Li, and Xu Tang · 2021
Cited alongside, same era.
Towards open world object detection
KJ Joseph, Salman Khan, Fahad Shahbaz Khan, and Vineeth N Balasubramanian · 2021
Cited alongside, same era.
Open world compositional zero-shot learning
Massimiliano Mancini, Muhammad Ferjad Naeem, Yongqin Xian, and Zeynep Akata · 2021
Cited alongside, same era.
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu · 2022
Later among the works it cites.
P3iv: Probabilistic procedure planning from instructional videos with weak supervision
He Zhao, Isma Hadji, Nikita Dvornik, Konstantinos G Derpanis, Richard P Wildes, and Allan D Jepson · 2022
Later among the works it cites.
Hiervl: Learning hierarchical video-language embeddings
Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani, and Kristen Grauman · 2023
Closest in time.
Opening the vocabulary of egocentric actions
Dibyadip Chatterjee, Fadime Sener, Shugao Ma, and Angela Yao · 2023
Closest in time.
Stepformer: Self-supervised step discovery and localization in instructional videos
Nikita Dvornik, Isma Hadji, Ran Zhang, Konstantinos G Derpanis, Richard P Wildes, and Allan D Jepson · 2023
Closest in time.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Closest in time.
Learning to ground instructional articles in videos through narrations
Effrosyni Mavroudi, Triantafyllos Afouras, and Lorenzo Torresani · 2023
Closest in time.
Gpt-4 technical report
OpenAI · 2023
Closest in time.
An outlook into the future of egocentric vision
Chiara Plizzari, Gabriele Goletto, Antonino Furnari, Siddhant Bansal, Francesco Ragusa, Giovanni Maria Farinella, Dima Damen, and Tatiana Tommasi · 2023
Closest in time.
Naq: Leveraging narrations as queries to supervise episodic memory
Santhosh Kumar Ramakrishnan, Ziad Al-Halah, and Kristen Grauman · 2023
Closest in time.
Chop & learn: Recognizing and generating object-state compositions
Nirat Saini, Hanyu Wang, Archana Swaminathan, Vinoj Jayasundara, Bo He, Kamal Gupta, and Abhinav Shrivastava · 2023
Closest in time.
Breaking the” object” in video object segmentation
Pavel Tokmakov, Jie Li, and Adrien Gaidon · 2023
Closest in time.
Localizing active objects from egocentric vision with symbolic world knowledge
Te-Lin Wu, Yu Zhou, and Nanyun Peng · 2023
Closest in time.
Video state-changing object segmentation
Jiangwei Yu, Xiang Li, Xinran Zhao, Hongming Zhang, and Yu-Xiong Wang · 2023
Closest in time.
Learning video representations from large language models
Yue Zhao, Ishan Misra, Philipp Krähenbühl, and Rohit Girdhar · 2023
Closest in time.