Fetching the paper…
Reading the bibliography…
Our goal is to learn a video representation that is useful for downstream procedure understanding tasks in instructional videos.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Unsupervised learning of video representations using lstms
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhudinov · 2015
Earlier work this paper cites.
Unsupervised learning from narrated instruction videos
Jean-Baptiste Alayrac, Piotr Bojanowski, Nishant Agrawal, Josef Sivic, Ivan Laptev, and Simon Lacoste-Julien · 2016
Earlier work this paper cites.
Deep predictive coding networks for video prediction and unsupervised learning
William Lotter, Gabriel Kreiman, and David Cox · 2016
Earlier work this paper cites.
Anticipating visual representations from unlabeled video
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Earlier work this paper cites.
Self-supervised video representation learning with odd-one-out networks
Basura Fernando, Hakan Bilen, Efstratios Gavves, and Stephen Gould · 2017
Earlier work this paper cites.
Unsupervised representation learning by sorting sequences
Hsin-Ying Lee, Jia-Bin Huang, Maneesh Singh, and Ming-Hsuan Yang · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Procnets: Learning to segment procedures in untrimmed and unconstrained videos
Luowei Zhou, Chenliang Xu, and Jason J Corso · 2017
Earlier work this paper cites.
Finding” it”: Weakly-supervised reference-aware visual grounding in instructional videos
De-An Huang, Shyamal Buch, Lucio Dery, Animesh Garg, Li Fei-Fei, and Juan Carlos Niebles · 2018
Earlier work this paper cites.
Wikihow: A large scale text summarization dataset
Mahnaz Koupaee and William Yang Wang · 2018
Earlier work this paper cites.
Towards automatic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason J Corso · 2018
Earlier work this paper cites.
Unsupervised procedure learning via joint dynamic summarization
Ehsan Elhamifar and Zwe Naing · 2019
Earlier work this paper cites.
Large-scale weakly-supervised pre-training for video action recognition
Deepti Ghadiyaram, Du Tran, and Dhruv Mahajan · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 2019
Earlier work this paper cites.
Self-supervised video representation learning with space-time cubic puzzles
Dahun Kim, Donghyeon Cho, and In So Kweon · 2019
Earlier work this paper cites.
Leveraging the present to anticipate the future in videos
Antoine Miech, Ivan Laptev, Josef Sivic, Heng Wang, Lorenzo Torresani, and Du Tran · 2019
Earlier work this paper cites.
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Learning video representations using contrastive bidirectional transformer
Chen Sun, Fabien Baradel, Kevin Murphy, and Cordelia Schmid · 2019
Earlier work this paper cites.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid · 2019
Earlier work this paper cites.
Coin: A large-scale dataset for comprehensive instructional video analysis
Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, and Jie Zhou · 2019
Earlier work this paper cites.
Self-supervised spatio-temporal representation learning for videos by predicting motion and appearance statistics
Jiangliu Wang, Jianbo Jiao, Linchao Bao, Shengfeng He, Yunhui Liu, and Wei Liu · 2019
Earlier work this paper cites.
Self-supervised spatiotemporal learning via video clip order prediction
Dejing Xu, Jun Xiao, Zhou Zhao, Jian Shao, Di Xie, and Yueting Zhuang · 2019
Earlier work this paper cites.
Cross-task weakly supervised learning from instructional videos
Dimitri Zhukov, Jean-Baptiste Alayrac, Ramazan Gokberk Cinbis, David Fouhey, Ivan Laptev, and Josef Sivic · 2019
Earlier work this paper cites.
Self-supervised multimodal versatile networks
Jean-Baptiste Alayrac, Adria Recasens, Rosalia Schneider, Relja Arandjelović, Jason Ramapuram, Jeffrey De Fauw, Lucas Smaira, Sander Dieleman, and Andrew Zisserman · 2020
Earlier work this paper cites.
Self-supervised learning by cross-modal audio-video clustering
Humam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani, Bernard Ghanem, and Du Tran · 2020
Earlier work this paper cites.
Speednet: Learning the speediness in videos
Sagie Benaim, Ariel Ephrat, Oran Lang, Inbar Mosseri, William T Freeman, Michael Rubinstein, Michal Irani, and Tali Dekel · 2020
Earlier work this paper cites.
Procedure planning in instructional videos
Chien-Yi Chang, De-An Huang, Danfei Xu, Ehsan Adeli, Li Fei-Fei, and Juan Carlos Niebles · 2020
Earlier work this paper cites.
Action modifiers: Learning from adverbs in instructional videos
Hazel Doughty, Ivan Laptev, Walterio Mayol-Cuevas, and Dima Damen · 2020
Earlier work this paper cites.
Self-supervised multi-task procedure learning from instructional videos
Ehsan Elhamifar and Dat Huynh · 2020
Cited alongside, same era.
Hero: Hierarchical encoder for video+ language omni-representation pre-training
Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu · 2020
Cited alongside, same era.
Univl: A unified video and language pre-training model for multimodal understanding and generation
Huaishao Luo, Lei Ji, Botian Shi, Haoyang Huang, Nan Duan, Tianrui Li, Jason Li, Taroon Bharti, and Ming Zhou · 2020
Cited alongside, same era.
End-to-end learning of visual representations from uncurated instructional videos
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman · 2020
Cited alongside, same era.
Mpnet: Masked and permuted pre-training for language understanding
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu · 2020
Cited alongside, same era.
Vlm: Task-agnostic video-language model pre-training for video understanding
Hu Xu, Gargi Ghosh, Po-Yao Huang, Prahal Arora, Masoumeh Aminzadeh, Christoph Feichtenhofer, Florian Metze, and Luke Zettlemoyer · 2021
Later among the works it cites.
Videoclip: Contrastive pre-training for zero-shot video-text understanding
Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer · 2021
Later among the works it cites.
Induce, edit, retrieve: Language grounded multimodal schema for instructional video retrieval
Yue Yang, Joongwon Kim, Artemis Panagopoulou, Mark Yatskar, and Chris Callison-Burch · 2021
Later among the works it cites.
Visual goal-step inference using wikihow
Yue Yang, Artemis Panagopoulou, Qing Lyu, Li Zhang, Mark Yatskar, and Chris Callison-Burch · 2021
Later among the works it cites.
Merlot: Multimodal neural script knowledge models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Comprehensive instructional video analysis: The coin dataset and performance evaluation
Yansong Tang, Jiwen Lu, and Jie Zhou · 2020
Cited alongside, same era.
Self-supervised video representation learning by pace prediction
Jiangliu Wang, Jianbo Jiao, and Yun-Hui Liu · 2020
Cited alongside, same era.
A benchmark for structured procedural knowledge extraction from cooking videos
Frank F Xu, Lei Ji, Botian Shi, Junyi Du, Graham Neubig, Yonatan Bisk, and Nan Duan · 2020
Cited alongside, same era.
Li Zhang, Qing Lyu, and Chris Callison-Burch · 2020
Cited alongside, same era.
Reasoning about goals, steps, and temporal ordering with wikihow
Li Zhang, Qing Lyu, and Chris Callison-Burch · 2020
Cited alongside, same era.
Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text
Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong · 2021
Cited alongside, same era.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Cited alongside, same era.
Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi · 2021
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, et al · 2022
Later among the works it cites.
Graph2vid: Flow graph to video grounding for weakly-supervised multi-step localization
Nikita Dvornik, Isma Hadji, Hai Pham, Dhaivat Bhatt, Brais Martinez, Afsaneh Fazly, and Allan D Jepson · 2022
Later among the works it cites.
Bridging video-text retrieval with multiple choice questions
Yuying Ge, Yixiao Ge, Xihui Liu, Dian Li, Ying Shan, Xiaohu Qie, and Ping Luo · 2022
Later among the works it cites.
Weakly-supervised online action segmentation in multi-view instructional videos
Reza Ghoddoosian, Isht Dwivedi, Nakul Agarwal, Chiho Choi, and Behzad Dariush · 2022
Later among the works it cites.
Hierarchical modeling for task recognition and action segmentation in weakly-labeled instructional videos
Reza Ghoddoosian, Saif Sayed, and Vassilis Athitsos · 2022
Later among the works it cites.
Learning temporal video procedure segmentation from an automatically collected large dataset
Lei Ji, Chenfei Wu, Daisy Zhou, Kun Yan, Edward Cui, Xilin Chen, and Nan Duan · 2022
Later among the works it cites.
Egotaskqa: Understanding human tasks in egocentric videos
Baoxiong Jia, Ting Lei, Song-Chun Zhu, and Siyuan Huang · 2022
Later among the works it cites.
Align and prompt: Video-and-language pre-training with entity prompts
Dongxu Li, Junnan Li, Hongdong Li, Juan Carlos Niebles, and Steven CH Hoi · 2022
Later among the works it cites.
Bridge-prompt: Towards ordinal action understanding in instructional videos
Muheng Li, Lei Chen, Yueqi Duan, Zhilan Hu, Jianjiang Feng, Jie Zhou, and Jiwen Lu · 2022
Later among the works it cites.
Order-constrained representation learning for instructional video prediction
Muheng Li, Lei Chen, Jiwen Lu, Jianjiang Feng, and Jie Zhou · 2022
Later among the works it cites.
Egocentric video-language pretraining
Kevin Qinghong Lin, Alex Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Zhongcong Xu, Difei Gao, Rongcheng Tu, Wenzhe Zhao, Weijie Kong, Chengfei Cai, Hongfa Wang, Dima Damen, Bernard Ghanem, Wei Liu, and Mike Zheng Shou · 2022
Later among the works it cites.
Learning to recognize procedural activities with distant supervision
Xudong Lin, Fabio Petroni, Gedas Bertasius, Marcus Rohrbach, Shih-Fu Chang, and Lorenzo Torresani · 2022
Later among the works it cites.
Set-supervised action learning in procedural task videos via pairwise order consistency
Zijia Lu and Ehsan Elhamifar · 2022
Later among the works it cites.
Survey: Transformer based video-language pre-training
Ludan Ruan and Qin Jin · 2022
Later among the works it cites.
Self-supervised learning for videos: A survey
Madeline C Schiappa, Yogesh S Rawat, and Mubarak Shah · 2022
Later among the works it cites.
End-to-end generative pretraining for multimodal video captioning
Paul Hongsuck Seo, Arsha Nagrani, Anurag Arnab, and Cordelia Schmid · 2022
Later among the works it cites.
Semi-weakly-supervised learning of complex actions from instructional task videos
Yuhan Shen and Ehsan Elhamifar · 2022
Later among the works it cites.
Look for the change: Learning object states and state-modifying actions from untrimmed web videos
Tomáš Souček, Jean-Baptiste Alayrac, Antoine Miech, Ivan Laptev, and Josef Sivic · 2022
Later among the works it cites.
Plate: Visually-grounded planning with transformers in procedural tasks
Jiankai Sun, De-An Huang, Bo Lu, Yun-Hui Liu, Bolei Zhou, and Animesh Garg · 2022
Later among the works it cites.
Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks
Yi-Lin Sung, Jaemin Cho, and Mohit Bansal · 2022
Later among the works it cites.
Multimedia generative script learning for task planning
Qingyun Wang, Manling Li, Hou Pong Chan, Lifu Huang, Julia Hockenmaier, Girish Chowdhary, and Heng Ji · 2022
Later among the works it cites.
Cross-modal contrastive distillation for instructional activity anticipation
Zhengyuan Yang, Jingen Liu, Jing Huang, Xiaodong He, Tao Mei, Chenliang Xu, and Jiebo Luo · 2022
Later among the works it cites.
Merlot reserve: Neural script knowledge through vision and language and sound
Rowan Zellers, Jiasen Lu, Ximing Lu, Youngjae Yu, Yanpeng Zhao, Mohammadreza Salehi, Aditya Kusupati, Jack Hessel, Ali Farhadi, and Yejin Choi · 2022
Later among the works it cites.
Show me more details: Discovering hierarchies of procedures from semi-structured web data
Shuyan Zhou, Li Zhang, Yue Yang, Qing Lyu, Pengcheng Yin, Chris Callison-Burch, and Graham Neubig · 2022
Later among the works it cites.