Fetching the paper…
Reading the bibliography…
Following step-by-step procedures is an essential component of various activities carried out by individuals in their daily lives.
Weakly Supervised Energy-Based Learning for Action Segmentation
Jun Li, Peng Lei, Peng Lei, and Sinisa Todorovic · 1909
Earlier work this paper cites.
Hierarchical learning in stochastic domains: Preliminary results
Leslie Pack Kaelbling · 1993
Earlier work this paper cites.
Hierarchical Language-based Representation of Events in Video Streams
Ram Nevatia, Tao Zhao, and Somboon Hongeng · 2003
Earlier work this paper cites.
Audio-Visual Instance Discrimination with Cross-Modal Agreement
Pedro Morgado, Nuno Vasconcelos, Ishan Misra, and Ishan Misra · 2004
Earlier work this paper cites.
The EPIC-KITCHENS Dataset: Collection, Challenges and Baselines
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, William Price, Will Price, Will Price, and Michael Wray · 2005
Earlier work this paper cites.
Learning to Segment Actions from Observation and Narration
Daniel Fried, Daniel Fried, Jean-Baptiste Alayrac, Phil Blunsom, Chris Dyer, Chris Dyer, Stephen Clark, and Aida Nematzadeh · 2005
Earlier work this paper cites.
MS-TCN: Multi-Stage Temporal Convolutional Network for Action Segmentation
Yazan Abu Farha, Juergen Gall, Jürgen Gall, and Juergen Gall · 2006
Earlier work this paper cites.
In the eye of the beholder: Gaze and actions in first person video
Yin Li, Miao Liu, and James M. Rehg · 2006
Earlier work this paper cites.
Guide to the carnegie mellon university multimodal activity (cmu-mmac) database
Fernando De la Torre, Jessica K. Hodgins, Adam W. Bargteil, Xavier Martin, J. Robert Macey, Alex Tusell Collado, and Pep Beltran · 2008
Earlier work this paper cites.
Development and evaluation of an ecological task to assess executive functioning post childhood tbi: The children’s cooking task
Mathilde P. Chevignard, Cathy Catroppa, Jane Galvin, and Vicki Anderson · 2010
Earlier work this paper cites.
Learning to recognize objects in egocentric activities
A. Fathi, Xiaofeng Ren, and J. M. Rehg · 2011
Earlier work this paper cites.
Learning to recognize objects in egocentric activities
Alireza Fathi, Xiaofeng Ren, and James M. Rehg · 2011
Earlier work this paper cites.
Parsing video events with goal inference and intent prediction
Mingtao Pei, Yunde Jia, and Song-Chun Zhu · 2011
Earlier work this paper cites.
Combining embedded accelerometers with computer vision for recognizing food preparation activities
Sebastian Stein and Stephen J. McKenna · 2013
Earlier work this paper cites.
Weakly supervised action labeling in videos under ordering constraints
Piotr Bojanowski, Rémi Lajugie, Francis Bach, Ivan Laptev, Jean Ponce, Cordelia Schmid, and Josef Sivic · 2014
Earlier work this paper cites.
THUMOS challenge: Action recognition with a large number of classes
Y.-G. Jiang, J. Liu, A. Roshan Zamir, G. Toderici, I. Laptev, M. Shah, and R. Sukthankar · 2014
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
Hilde Kuehne, Ali Arslan, and Thomas Serre · 2014
Earlier work this paper cites.
The Language of Actions: Recovering the Syntax and Semantics of Goal-Directed Human Activities
Hilde Kuehne, Ali Bilgin Arslan, and Thomas Serre · 2014
Earlier work this paper cites.
Weakly-Supervised Alignment of Video With Text
Piotr Bojanowski, Rémi Lajugie, Edouard Grave, Francis Bach, Ivan Laptev, Jean Ponce, Cordelia Schmid, and Cordelia Schmid · 2015
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Bernard Ghanem Fabian Caba Heilbron, Victor Escorcia and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
What’s Cookin’? Interpreting Cooking Videos using Text, Speech and Vision
Jonathan Malmaud, Jonathan Huang, Vivek Rathod, Nick Johnston, Andrew Rabinovich, Kevin Murphy, Kevin Murphy, Kevin Murphy, and Kevin Murphy · 2015
Earlier work this paper cites.
Recognizing fine-grained and composite activities using hand-centric features and script data
Marcus Rohrbach, Anna Rohrbach, Michaela Regneri, Sikandar Amin, Mykhaylo Andriluka, Manfred Pinkal, and Bernt Schiele · 2015
Earlier work this paper cites.
Generating notifications for missing actions: Don’t forget to turn the lights off!
Bilge Soran, Ali Farhadi, and Linda Shapiro · 2015
Earlier work this paper cites.
Autonomously constructing hierarchical task networks for planning and human-robot collaboration
Bradley Hayes and Brian Scassellati · 2016
Earlier work this paper cites.
Connectionist temporal modeling for weakly supervised action labeling
De-An Huang, Li Fei-Fei, and Juan Carlos Niebles · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A. Sigurdsson, Gül Varol, X. Wang, Ali Farhadi, Ivan Laptev, and Abhinav Kumar Gupta · 2016
Earlier work this paper cites.
Ava: A video dataset of spatio-temporally localized atomic visual actions
Chunhui Gu, Chen Sun, Sudheendra Vijayanarasimhan, Caroline Pantofaru, David A. Ross, George Toderici, Yeqing Li, Susanna Ricco, Rahul Sukthankar, Cordelia Schmid, and Jitendra Malik · 2017
Earlier work this paper cites.
Learning Spatio-Temporal Features with 3D Residual Networks for Action Recognition, August 2017
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh · 2017
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization, January 2017
Diederik P. Kingma and Jimmy Ba · 2017
Earlier work this paper cites.
Online real-time multiple spatiotemporal action localisation and prediction
Gurkirt Singh, Suman Saha, Michael Sapienza, Philip H. S. Torr, and Fabio Cuzzolin · 2017
Earlier work this paper cites.
Towards Automatic Learning of Procedures from Web Instructional Videos
Luowei Zhou, Chenliang Xu, and Jason J. Corso · 2017
Earlier work this paper cites.
Localizing Moments in Video with Temporal Language
Lisa Anne Hendricks, Bryan C. Russell, Eli Shechtman, Josef Sivic, Oliver Wang, and Trevor Darrell · 2018
Earlier work this paper cites.
Towards automatic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason J. Corso · 2018
Earlier work this paper cites.
D3TW: discriminative differentiable dynamic time warping for weakly supervised action alignment and segmentation
Chien-Yi Chang, De-An Huang, Yanan Sui, Li Fei-Fei, and Juan Carlos Niebles · 2019
Earlier work this paper cites.
Temporal cycle-consistency learning
Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson, Pierre Sermanet, and Andrew Zisserman · 2019
Earlier work this paper cites.
Unsupervised procedure learning via joint dynamic summarization
Ehsan Elhamifar and Zwe Naing · 2019
Earlier work this paper cites.
SlowFast Networks for Video Recognition, October 2019
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Earlier work this paper cites.
Epic-tent: An egocentric video dataset for camping tent assembly
Youngkyoon Jang, Brian Sullivan, Casimir Ludwig, Iain Gilchrist, Dima Damen, and Walterio Mayol-Cuevas · 2019
Earlier work this paper cites.
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Cited alongside, same era.
Zero-shot anticipation for instructional activities
Fadime Sener and Angela Yao · 2019
Cited alongside, same era.
COIN: A Large-Scale Dataset for Comprehensive Instructional Video Analysis
Yansong Tang, Dajun Ding, Dajun Ding, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, and Jie Zhou · 2019
Cited alongside, same era.
Cross-task weakly supervised learning from instructional videos
Dimitri Zhukov, Jean-Baptiste Alayrac, Ramazan Gokberk Cinbis, David F. Fouhey, Ivan Laptev, and Josef Sivic · 2019
Cited alongside, same era.
Robust learning of tractable probabilistic models
Rohith Peddi, Tahrima Rahman, and Vibhav Gogate · 2022
Later among the works it cites.
MAViL: Masked Audio-Video Learners
Po-Yao Huang, Vasu Sharma, Hu Xu, Chaitanya K. Ryali, Haoqi Fan, Yanghao Li, Shang-Wen Li, Gargi Ghosh, J. Malik, and Christoph Feichtenhofer · 2022
Later among the works it cites.
SVIP: sequence verification for procedures in videos
Yicheng Qian, Weixin Luo, Dongze Lian, Xu Tang, Peilin Zhao, and Shenghua Gao · 2022
Later among the works it cites.
Self-Supervised Predictive Convolutional Attentive Block for Anomaly Detection, March 2022
Nicolae-Catalin Ristea, Neelu Madan, Radu Tudor Ionescu, Kamal Nasrollahi, Fahad Shahbaz Khan, Thomas B. Moeslund, and Mubarak Shah · 2022
Later among the works it cites.
XKD: Cross-modal Knowledge Distillation with Domain Alignment for Video Representation Learning
Sarkar, Pritam and Etemad, Ali · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Self-supervised multi-task procedure learning from instructional videos
Ehsan Elhamifar and Dat Huynh · 2020
Cited alongside, same era.
X3D: Expanding Architectures for Efficient Video Recognition, April 2020
Christoph Feichtenhofer · 2020
Cited alongside, same era.
Daily performance of adolescents with executive function deficits: An empirical study using a complex-cooking task
Yael Fogel, Sara Rosenblum, Renana Hirsh, Mathilde Chevignard, and Naomi Josman · 2020
Cited alongside, same era.
Adversarial generative grammars for human activity prediction
A. J. Piergiovanni, Anelia Angelova, Alexander Toshev, and Michael S. Ryoo · 2020
Cited alongside, same era.
English recipe flow graph corpus
Yoko Yamakata, Shinsuke Mori, and John Carroll · 2020
Cited alongside, same era.
Learning Actionness via Long-Range Temporal Order Verification
Dimitri Zhukov, Dimitri Zhukov, Jean-Baptiste Alayrac, Ivan Laptev, and Josef Sivic · 2020
Cited alongside, same era.
Visual Semantic Role Labeling for Video Understanding
Arka Sadhu, Tanmay Gupta, Mark Yatskar, R. Nevatia, and Aniruddha Kembhavi · 2021
Cited alongside, same era.
Multi-Task Learning of Object State Changes from Uncurated Videos
Souček, Tomáš, Alayrac, Jean-Baptiste, Miech, Antoine, Laptev, Ivan, and Sivic, Josef · 2022
Later among the works it cites.
Long-Form Video-Language Pre-Training with Multimodal Temporal Contrastive Learning
Sun, Yuchong, Xue, Hongwei, Song, Ruihua, Liu, Bei, Yang, Huan, and Fu, Jianlong · 2022
Later among the works it cites.
Temporal Alignment Networks for Long-term Video
Tengda Han, Weidi Xie, and Andrew Zisserman · 2022
Later among the works it cites.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang · 2022
Later among the works it cites.
Learning to Align Sequential Actions in the Wild
Weizhe Liu, Bugra Tekin, Huseyin Coskun, Vibhav Vineet, P. Fua, and M. Pollefeys · 2022
Later among the works it cites.
SVIP: Sequence VerIfication for Procedures in Videos
Yichen Qian, Weixin Luo, Dongze Lian, Xu Tang, P. Zhao, and Shenghua Gao · 2022
Later among the works it cites.
Learning Video Representations from Large Language Models
Yue Zhao, Ishan Misra, Philipp Krahenbuhl, and Rohit Girdhar · 2022
Later among the works it cites.
Actionformer: Localizing moments of actions with transformers
Chen-Lin Zhang, Jianxin Wu, and Yin Li · 2022
Later among the works it cites.
Collaborative Reasoning on Multi-Modal Semantic Graphs for Video-Grounded Dialogue Generation
Zhao, Xueliang, Wang, Yuxuan, Tao, Chongyang, Wang, Chenshuo, and Zhao, Dongyan · 2022
Later among the works it cites.
Learning Action Changes by Measuring Verb-Adverb Textual Relationships
D. Moltisanti, Frank Keller, Hakan Bilen, and Laura Sevilla-Lara · 2023
Closest in time.
Weakly-supervised online action segmentation in multi-view instructional videos
Reza Ghoddoosian, Isht Dwivedi, Nakul Agarwal, Chiho Choi, and Behzad Dariush · 2023
Closest in time.
Imagebind: One embedding space to bind them all
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra · 2023
Closest in time.
Egotv: Egocentric task verification from natural language task descriptions
Rishi Hazra, Brian Chen, Akshara Rai, Nitin Kamra, and Ruta Desai · 2023
Closest in time.
Procedure-Aware Pretraining for Instructional Video Understanding
Honglu Zhou, Roberto Mart’in-Mart’in, M. Kapadia, S. Savarese, and Juan Carlos Niebles · 2023
Closest in time.
Aligning Step-by-Step Instructional Diagrams to Video Demonstrations
Jiahao Zhang, A. Cherian, Yanbin Liu, Yizhak Ben-Shabat, Cristián Rodríguez, and Stephen Gould · 2023
Closest in time.
Unsupervised Task Graph Generation from Instructional Video Transcripts
L. Logeswaran, Sungryull Sohn, Y. Jang, Moontae Lee, and Ho Hin Lee · 2023
Closest in time.
Video-llava: Learning united visual representation by alignment before projection
Bin Lin, Bin Zhu, Yang Ye, Munan Ning, Peng Jin, and Li Yuan · 2023
Closest in time.
Action dynamics task graphs for learning plannable representations of procedural tasks, 2023
Weichao Mao, Ruta Desai, Michael Louis Iuzzolino, and Nitin Kamra · 2023
Closest in time.
StepFormer: Self-supervised Step Discovery and Localization in Instructional Videos
Nikita Dvornik, Isma Hadji, Ran Zhang, K. Derpanis, Animesh Garg, R. Wildes, and A. Jepson · 2023
Closest in time.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever · 2023
Closest in time.
Timechat: A time-sensitive multimodal large language model for long video understanding
Shuhuai Ren, Linli Yao, Shicheng Li, Xu Sun, and Lu Hou · 2023
Closest in time.
Action scene graphs for long-form understanding of egocentric videos, 2023
Ivan Rodin, Antonino Furnari, Kyle Min, Subarna Tripathi, and Giovanni Maria Farinella · 2023
Closest in time.
VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset
Sihan Chen, Xingjian He, Longteng Guo, Xinxin Zhu, Weining Wang, Jinhui Tang, and Jing Liu · 2023
Closest in time.
Weakly Supervised Video Representation Learning with Unaligned Text for Sequential Videos
Sixun Dong, Huazhang Hu, Dongze Lian, Weixin Luo, Yichen Qian, and Shenghua Gao · 2023
Closest in time.
Llama: Open and efficient foundation language models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Closest in time.
Holoassist: an egocentric human interaction dataset for interactive ai assistants in the real world
Xin Wang, Taein Kwon, Mahdi Rad, Bowen Pan, Ishani Chakraborty, Sean Andrist, Dan Bohus, Ashley Feniello, Bugra Tekin, Felipe Vieira Frujeri, Neel Joshi, and Marc Pollefeys · 2023
Closest in time.
Visual Transformation Telling
Xin Hong, Yanyan Lan, Liang Pang, J. Guo, and Xueqi Cheng · 2023
Closest in time.
Tagging before Alignment: Integrating Multi-Modal Tags for Video-Text Retrieval
Yizhen Chen, Jie Wang, Lijian Lin, Zhongang Qi, Jin Ma, and Ying Shan · 2023
Closest in time.
Prego: online mistake detection in procedural egocentric videos, 2024
Alessandro Flaborea, Guido Maria D’Amely di Melendugno, Leonardo Plini, Luca Scofano, Edoardo De Matteis, Antonino Furnari, Giovanni Maria Farinella, and Fabio Galasso · 2024
Closest in time.
Towards unbiased and robust spatio-temporal scene graph generation and anticipation, 2024
Rohith Peddi, Saurabh, Ayush Abhay Shrivastava, Parag Singla, and Vibhav Gogate · 2024
Closest in time.
IndustReal: A dataset for procedure step recognition handling execution errors in egocentric videos in an industrial-like setting, 2024
Tim J. Schoonbeek, Tim Houben, Hans Onvlee, Peter H. N. de With, and Fons van der Sommen · 2024
Closest in time.
Differentiable task graph learning: Procedural activity representation and online mistake detection from egocentric videos, 2024
Luigi Seminara, Giovanni Maria Farinella, and Antonino Furnari · 2024
Closest in time.
Towards scene graph anticipation
Rohith Peddi, Saksham Singh, Saurabh, Parag Singla, and Vibhav Gogate · 2025
Closest in time.