Fetching the paper…
Reading the bibliography…
In this paper we present an approach for localizing steps of procedural activities in narrated how-to videos.
Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks
Dong-Hyun Lee et al · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean · 2013
Earlier work this paper cites.
Weakly supervised action labeling in videos under ordering constraints
Piotr Bojanowski, Rémi Lajugie, Francis Bach, Ivan Laptev, Jean Ponce, Cordelia Schmid, and Josef Sivic · 2014
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
Hilde Kuehne, Ali Arslan, and Thomas Serre · 2014
Earlier work this paper cites.
Weakly-supervised alignment of video with text
Piotr Bojanowski, Rémi Lajugie, Edouard Grave, Francis Bach, Ivan Laptev, Jean Ponce, and Cordelia Schmid · 2015
Earlier work this paper cites.
What’s cookin’? interpreting cooking videos using text, speech and vision
Jonathan Malmaud, Jonathan Huang, Vivek Rathod, Nicholas Johnston, Andrew Rabinovich, and Kevin Murphy · 2015
Earlier work this paper cites.
Unsupervised semantic parsing of video collections
Ozan Sener, Amir R Zamir, Silvio Savarese, and Ashutosh Saxena · 2015
Earlier work this paper cites.
Unsupervised learning from narrated instruction videos
Jean-Baptiste Alayrac, Piotr Bojanowski, Nishant Agrawal, Josef Sivic, Ivan Laptev, and Simon Lacoste-Julien · 2016
Earlier work this paper cites.
An end-to-end generative framework for video segmentation and recognition
Hilde Kuehne, Juergen Gall, and Thomas Serre · 2016
Earlier work this paper cites.
Temporal convolutional networks: A unified approach to action segmentation
Colin Lea, Rene Vidal, Austin Reiter, and Gregory D Hager · 2016
Earlier work this paper cites.
A multi-stream bi-directional recurrent neural network for fine-grained action detection
Bharat Singh, Tim K Marks, Michael Jones, Oncel Tuzel, and Ming Shao · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Earlier work this paper cites.
Weakly supervised learning of actions from transcripts
Hilde Kuehne, Alexander Richard, and Juergen Gall · 2017
Earlier work this paper cites.
Wikihow: A large scale text summarization dataset
Mahnaz Koupaee and William Yang Wang · 2018
Earlier work this paper cites.
Action sets: Weakly supervised action segmentation without ordering constraints
Alexander Richard, Hilde Kuehne, and Juergen Gall · 2018
Earlier work this paper cites.
Neuralnetwork-viterbi: A framework for weakly supervised video learning
Alexander Richard, Hilde Kuehne, Ahsan Iqbal, and Juergen Gall · 2018
Earlier work this paper cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy · 2018
Earlier work this paper cites.
D3tw: Discriminative differentiable dynamic time warping for weakly supervised action alignment and segmentation
C. Chang, De-An Huang, Yanan Sui, Li Fei-Fei, and Juan Carlos Niebles · 2019
Earlier work this paper cites.
Unicoder: A universal language encoder by pre-training with multiple cross-lingual tasks
Haoyang Huang, Yaobo Liang, Nan Duan, Ming Gong, Linjun Shou, Daxin Jiang, and Ming Zhou · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Ivan Laptev, and Josef Sivic · 2019
Earlier work this paper cites.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin P. Murphy, and Cordelia Schmid · 2019
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal · 2019
Earlier work this paper cites.
Coin: A large-scale dataset for comprehensive instructional video analysis
Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, and Jie Zhou · 2019
Earlier work this paper cites.
Multimodal transformer for unaligned multimodal language sequences
Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J. Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
Cross-task weakly supervised learning from instructional videos
Dimitri Zhukov, Jean-Baptiste Alayrac, Ramazan Gokberk Cinbis, David Fouhey, Ivan Laptev, and Josef Sivic · 2019
Cited alongside, same era.
Few-shot video classification via temporal alignment
Kaidi Cao, Jingwei Ji, Zhangjie Cao, Chien-Yi Chang, and Juan Carlos Niebles · 2020
Cited alongside, same era.
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2020
Cited alongside, same era.
Self-supervised multi-task procedure learning from instructional videos
Ehsan Elhamifar and Dat Huynh · 2020
Cited alongside, same era.
Learning to segment actions from observation and narration
VideoCLIP: Contrastive pre-training for zero-shot video-text understanding
Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer · 2021
Later among the works it cites.
Taco: Token-aware cascade contrastive learning for video-text alignment
Jianwei Yang, Yonatan Bisk, and Jianfeng Gao · 2021
Later among the works it cites.
Asformer: Transformer for action segmentation
Fangqiu Yi, Hongyu Wen, and Tingting Jiang · 2021
Later among the works it cites.
Locvtp: Video-text pre-training for temporal localization
Meng Cao, Tianyu Yang, Junwu Weng, Can Zhang, Jue Wang, and Yuexian Zou · 2022
Later among the works it cites.
Weakly-supervised temporal article grounding
Long Chen, Yulei Niu, Brian Chen, Xudong Lin, Guangxing Han, Christopher Thomas, Hammad Ayyubi, Heng Ji, and Shih-Fu Chang · 2022
Later among the works it cites.
Flow graph to video grounding for weakly-supervised multi-step localization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daniel Fried, Jean-Baptiste Alayrac, Phil Blunsom, Chris Dyer, Stephen Clark, and Aida Nematzadeh · 2020
Cited alongside, same era.
Coot: Cooperative hierarchical transformer for video-text representation learning
Simon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, and Thomas Brox · 2020
Cited alongside, same era.
HERO: Hierarchical encoder for Video+Language omni-representation pre-training
Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu · 2020
Cited alongside, same era.
Univl: A unified video and language pre-training model for multimodal understanding and generation
Huaishao Luo, Lei Ji, Botian Shi, Haoyang Huang, Nan Duan, Tianrui Li, Jason Li, Taroon Bharti, and Ming Zhou · 2020
Cited alongside, same era.
End-to-End Learning of Visual Representations from Uncurated Instructional Videos
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman · 2020
Cited alongside, same era.
Avlnet: Learning audio-visual language representations from instructional videos
Andrew Rouditchenko, Angie Boggust, David Harwath, Brian Chen, Dhiraj Joshi, Samuel Thomas, Kartik Audhkhasi, Hilde Kuehne, Rameswar Panda, Rogerio Feris, et al · 2020
Cited alongside, same era.
Fixmatch: Simplifying semi-supervised learning with consistency and confidence
Kihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang, Han Zhang, Colin A Raffel, Ekin Dogus Cubuk, Alexey Kurakin, and Chun-Liang Li · 2020
Cited alongside, same era.
Nikita Dvornik, Isma Hadji, Hai Pham, Dhaivat Bhatt, Brais Martinez, Afsaneh Fazly, and Allan D Jepson · 2022
Later among the works it cites.
Temporal alignment networks for long-term video
Tengda Han, Weidi Xie, and Andrew Zisserman · 2022
Later among the works it cites.
Video-text representation learning via differentiable weak temporal alignment
Dohwan Ko, Joonmyung Choi, Juyeon Ko, Shinyeong Noh, Kyoung-Woon On, Eun-Sol Kim, and Hyunwoo J Kim · 2022
Later among the works it cites.
Learning to recognize procedural activities with distant supervision
Xudong Lin, Fabio Petroni, Gedas Bertasius, Marcus Rohrbach, Shih-Fu Chang, and Lorenzo Torresani · 2022
Later among the works it cites.
X-CLIP:: End-to-end multi-grained contrastive learning for video-text retrieval
Yiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan, Ji Zhang, and Rongrong Ji · 2022
Later among the works it cites.
Simvtp: Simple video text pre-training with masked autoencoders
Yue Ma, Tianyu Yang, Yin Shan, and Xiu Li · 2022
Later among the works it cites.
Normalized contrastive learning for text-video retrieval
Yookoon Park, Mahmoud Azab, Seungwhan Moon, Bo Xiong, Florian Metze, Gourab Kundu, and Kirmani Ahmed · 2022
Later among the works it cites.
Semi-weakly-supervised learning of complex actions from instructional task videos
Yuhan Shen and Ehsan Elhamifar · 2022
Later among the works it cites.
Everything at once-multi-modal fusion transformer for video retrieval
Nina Shvetsova, Brian Chen, Andrew Rouditchenko, Samuel Thomas, Brian Kingsbury, Rogerio S Feris, David Harwath, James Glass, and Hilde Kuehne · 2022
Later among the works it cites.
VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang · 2022
Later among the works it cites.
Semi-supervised semantic segmentation using unreliable pseudo-labels
Yuchao Wang, Haochen Wang, Yujun Shen, Jingjing Fei, Wei Li, Guoqiang Jin, Liwei Wu, Rui Zhao, and Xinyi Le · 2022
Later among the works it cites.
Multimodal learning with transformers: A survey, 2022
Peng Xu, Xiatian Zhu, and David A. Clifton · 2022
Later among the works it cites.
Temporal alignment representation with contrastive learning
Yuncong Yang, Jiawei Ma, Shiyuan Huang, Long Chen, Xudong Lin, Guangxing Han, and Shih-Fu Chang · 2022
Later among the works it cites.
P3iv: Probabilistic procedure planning from instructional videos with weak supervision
He Zhao, Isma Hadji, Nikita Dvornik, Konstantinos G Derpanis, Richard P Wildes, and Allan D Jepson · 2022
Later among the works it cites.
Tempclr: Reconstructing hands via time-coherent contrastive learning
Andrea Ziani, Zicong Fan, Muhammed Kocabas, Sammy Christen, and Otmar Hilliges · 2022
Later among the works it cites.
Ego-only: Egocentric action detection without exocentric pretraining, 2023
Huiyu Wang, Mitesh Kumar Singh, and Lorenzo Torresani · 2023
Closest in time.
Pdpp: Projected diffusion for procedure planning in instructional videos
Hanlin Wang, Yilu Wu, Sheng Guo, and Limin Wang · 2023
Closest in time.
Learning procedure-aware video representation from instructional videos and their narrations
Yiwu Zhong, Licheng Yu, Yang Bai, Shangwen Li, Xueting Yan, and Yin Li · 2023
Closest in time.