Fetching the paper…
Reading the bibliography…
Video-language pre-training (VLP) has become increasingly important due to its ability to generalize to various vision and language tasks.
Discovering important people and objects for egocentric video summarization
Yong Jae Lee, Joydeep Ghosh, and Kristen Grauman · 2012
Earlier work this paper cites.
Sentibank: large-scale ontology and classifiers for detecting sentiment and emotions in visual content
Damian Borth, Tao Chen, Rongrong Ji, and Shih-Fu Chang · 2013
Earlier work this paper cites.
Large-scale video summarization using web-image priors
Aditya Khosla, Raffay Hamid, Chih-Jen Lin, and Neel Sundaresan · 2013
Earlier work this paper cites.
Diverse sequential subset selection for supervised video summarization
Boqing Gong, Wei-Lun Chao, Kristen Grauman, and Fei Sha · 2014
Earlier work this paper cites.
Joint summarization of large-scale collections of web images and videos for storyline reconstruction
Gunhee Kim, Leonid Sigal, and Eric P Xing · 2014
Earlier work this paper cites.
Category-specific video summarization
Danila Potapov, Matthijs Douze, Zaid Harchaoui, and Cordelia Schmid · 2014
Earlier work this paper cites.
Videoset: Video summary evaluation through text
Serena Yeung, Alireza Fathi, and Li Fei-Fei · 2014
Earlier work this paper cites.
Quasi real-time summarization for consumer videos
Bin Zhao and Eric P Xing · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Video co-summarization: Video summarization by visual co-occurrence
Wen-Sheng Chu, Yale Song, and Alejandro Jaimes · 2015
Earlier work this paper cites.
Tagging personal photos with transfer deep learning
Jianlong Fu, Tao Mei, Kuiyuan Yang, Hanqing Lu, and Yong Rui · 2015
Earlier work this paper cites.
Delving into egocentric actions
Yin Li, Zhefan Ye, and James M Rehg · 2015
Earlier work this paper cites.
Tvsum: Summarizing web videos using titles
Yale Song, Jordi Vallmitjana, Amanda Stent, and Alejandro Jaimes · 2015
Earlier work this paper cites.
Gaze-enabled egocentric video summarization via constrained submodular maximization
Jia Xu, Lopamudra Mukherjee, Yin Li, Jamieson Warner, James M Rehg, and Vikas Singh · 2015
Earlier work this paper cites.
Query-focused extractive video summarization
Aidean Sharghi, Boqing Gong, and Mubarak Shah · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
Summary transfer: Exemplar-based subset selection for video summarization
Ke Zhang, Wei-Lun Chao, Fei Sha, and Kristen Grauman · 2016
Earlier work this paper cites.
Video summarization with long short-term memory
Ke Zhang, Wei-Lun Chao, Fei Sha, and Kristen Grauman · 2016
Earlier work this paper cites.
Tall: Temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Earlier work this paper cites.
Query-focused video summarization: Dataset, evaluation, and a memory network based approach
Aidean Sharghi, Jacob S Laurel, and Boqing Gong · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang · 2017
Earlier work this paper cites.
Localizing moments in video with temporal language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
Actor and observer: Joint modeling of first and third-person videos
Gunnar A Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari · 2018
Earlier work this paper cites.
Charades-ego: A large-scale dataset of paired third and first person videos
Gunnar A Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari · 2018
Earlier work this paper cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy · 2018
Earlier work this paper cites.
A joint sequence fusion model for video question answering and retrieval
Youngjae Yu, Jongseok Kim, and Gunhee Kim · 2018
Earlier work this paper cites.
Query-conditioned three-player adversarial network for video summarization
Yujia Zhang, Michael Kampffmeyer, Xiaodan Liang, Min Tan, and Eric P Xing · 2018
Earlier work this paper cites.
Towards automatic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason Corso · 2018
Earlier work this paper cites.
Neural storyboard artist: Visualizing stories with coherent image sequences
Shizhe Chen, Bei Liu, Jianlong Fu, Ruihua Song, Qin Jin, Pingping Lin, Xiaoyu Qi, Chunting Wang, and Jin Zhou · 2019
Earlier work this paper cites.
Cross-lingual language model pretraining
Alexis Conneau and Guillaume Lample · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Summarizing videos with attention
Jiri Fajtl, Hajar Sadeghi Sokeh, Vasileios Argyriou, Dorothy Monekosso, and Paolo Remagnino · 2019
Earlier work this paper cites.
Heterogeneous memory enhanced multimodal attention model for video question answering
Chenyou Fan, Xiaofan Zhang, Shu Zhang, Wensheng Wang, Chi Zhang, and Heng Huang · 2019
Earlier work this paper cites.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Earlier work this paper cites.
Hierarchical variational network for user-diversified & query-focused video summarization
Pin Jiang and Yahong Han · 2019
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 2019
Earlier work this paper cites.
Emotion reinforced visual storytelling
Nanxing Li, Bei Liu, Zhizhong Han, Yu-Shen Liu, and Jianlong Fu · 2019
Earlier work this paper cites.
Beyond rnns: Positional self-attention with co-attention for video question answering
Xiangpeng Li, Jingkuan Song, Lianli Gao, Xianglong Liu, Wenbing Huang, Xiangnan He, and Chuang Gan · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Cited alongside, same era.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Vlg-net: Video-language graph matching network for video grounding
Mattia Soldan, Mengmeng Xu, Sisi Qu, Jesper Tegner, and Bernard Ghanem · 2021
Later among the works it cites.
Vlm: Task-agnostic video-language model pre-training for video understanding
Hu Xu, Gargi Ghosh, Po-Yao Huang, Prahal Arora, Masoumeh Aminzadeh, Christoph Feichtenhofer, Florian Metze, and Luke Zettlemoyer · 2021
Later among the works it cites.
Videoclip: Contrastive pre-training for zero-shot video-text understanding
Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer · 2021
Later among the works it cites.
Taco: Token-aware cascade contrastive learning for video-text alignment
Jianwei Yang, Yonatan Bisk, and Jianfeng Gao · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid · 2019
Cited alongside, same era.
Fine-grained action retrieval through multiple parts-of-speech embeddings
Michael Wray, Diane Larlus, Gabriela Csurka, and Dima Damen · 2019
Cited alongside, same era.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le · 2019
Cited alongside, same era.
Rethinking zero-shot video classification: End-to-end training for realistic applications
Biagio Brattoli, Joseph Tighe, Fedor Zhdanov, Pietro Perona, and Krzysztof Chalupka · 2020
Cited alongside, same era.
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2020
Cited alongside, same era.
Unsupervised and semi-supervised domain adaptation for action recognition from drones
Jinwoo Choi, Gaurav Sharma, Manmohan Chandraker, and Jia-Bin Huang · 2020
Cited alongside, same era.
Lemma: A multi-view dataset for le arning m ulti-agent m ulti-task a ctivities
Baoxiong Jia, Yixin Chen, Siyuan Huang, Yixin Zhu, and Song-chun Zhu · 2020
Cited alongside, same era.
Merlot: Multimodal neural script knowledge models
Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi · 2021
Later among the works it cites.
Video self-stitching graph network for temporal action localization
Chen Zhao, Ali K Thabet, and Bernard Ghanem · 2021
Later among the works it cites.
Vlmo: Unified vision-language pre-training with mixture-of-modality-experts
Hangbo Bao, Wenhui Wang, Li Dong, Qiang Liu, Owais Khan Mohammed, Kriti Aggarwal, Subhojit Som, Songhao Piao, and Furu Wei · 2022
Later among the works it cites.
Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al · 2022
Later among the works it cites.
Coarse-to-fine vision-language pre-training with fusion in the backbone
Zi-Yi Dou, Aishwarya Kamath, Zhe Gan, Pengchuan Zhang, Jianfeng Wang, Linjie Li, Zicheng Liu, Ce Liu, Yann LeCun, Nanyun Peng, et al · 2022
Later among the works it cites.
An empirical study of training end-to-end vision-and-language transformers
Zi-Yi Dou, Yichong Xu, Zhe Gan, Jianfeng Wang, Shuohang Wang, Lijuan Wang, Chenguang Zhu, Pengchuan Zhang, Lu Yuan, Nanyun Peng, et al · 2022
Later among the works it cites.
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al · 2022
Later among the works it cites.
Egotaskqa: Understanding human tasks in egocentric videos
Baoxiong Jia, Ting Lei, Song-Chun Zhu, and Siyuan Huang · 2022
Later among the works it cites.
Lavender: Unifying video-language understanding as masked language modeling
Linjie Li, Zhe Gan, Kevin Lin, Chung-Ching Lin, Zicheng Liu, Ce Liu, and Lijuan Wang · 2022
Later among the works it cites.
Cross-modal representation learning for zero-shot action recognition
Chung-Ching Lin, Kevin Lin, Lijuan Wang, Zicheng Liu, and Linjie Li · 2022
Later among the works it cites.
Egocentric video-language pretraining
Kevin Qinghong Lin, Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Zhongcong Xu, Denial Gao, Rong-Cheng Tu, Wenzhe Zhao, Weijie Kong, et al · 2022
Later among the works it cites.
Video swin transformer
Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu · 2022
Later among the works it cites.
Imu2clip: Multimodal contrastive learning for imu motion sensors from egocentric videos and text
Seungwhan Moon, Andrea Madotto, Zhaojiang Lin, Alireza Dirafzoon, Aparajita Saraf, Amy Bearman, and Babak Damavandi · 2022
Later among the works it cites.
Volta: Vision-language transformer with weakly-supervised local-feature alignment
Shraman Pramanick, Li Jing, Sayan Nag, Jiachen Zhu, Hardik Shah, Yann LeCun, and Rama Chellappa · 2022
Later among the works it cites.
Multimodal learning using optimal transport for sarcasm and humor detection
Shraman Pramanick, Aniket Roy, and Vishal M Patel · 2022
Later among the works it cites.
Long-form video-language pre-training with multimodal temporal contrastive learning
Yuchong Sun, Hongwei Xue, Ruihua Song, Bei Liu, Huan Yang, and Jianlong Fu · 2022
Later among the works it cites.
Object-aware video-language pre-training for retrieval
Jinpeng Wang, Yixiao Ge, Guanyu Cai, Rui Yan, Xudong Lin, Ying Shan, Xiaohu Qie, and Mike Zheng Shou · 2022
Later among the works it cites.
Image as a foreign language: Beit pretraining for all vision and vision-language tasks
Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, et al · 2022
Later among the works it cites.
Intentvizor: Towards generic query guided interactive video summarization
Guande Wu, Jianzhe Lin, and Claudio T Silva · 2022
Later among the works it cites.
Advancing high-resolution video-language representation with large-scale video transcriptions
Hongwei Xue, Tiankai Hang, Yanhong Zeng, Yuchong Sun, Bei Liu, Huan Yang, Jianlong Fu, and Baining Guo · 2022
Later among the works it cites.
Unified contrastive learning in image-text-label space
Jianwei Yang, Chunyuan Li, Pengchuan Zhang, Bin Xiao, Ce Liu, Lu Yuan, and Jianfeng Gao · 2022
Later among the works it cites.
Unitab: Unifying text and box outputs for grounded vision-language modeling
Zhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu, Faisal Ahmed, Zicheng Liu, Yumao Lu, and Lijuan Wang · 2022
Later among the works it cites.
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu · 2022
Later among the works it cites.
Lit: Zero-shot transfer with locked-image text tuning
Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, and Lucas Beyer · 2022
Later among the works it cites.
Hiervl: Learning hierarchical video-language embeddings
Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani, and Kristen Grauman · 2023
Closest in time.
Unifying vision-language representation space with single-tower transformer
Jiho Jang, Chaerin Kong, Donghyeon Jeon, Seonhoon Kim, and Nojun Kwak · 2023
Closest in time.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Closest in time.
Scaling language-image pre-training via masking
Yanghao Li, Haoqi Fan, Ronghang Hu, Christoph Feichtenhofer, and Kaiming He · 2023
Closest in time.
Univtg: Towards unified video-language temporal grounding, 2023
Kevin Qinghong Lin, Pengchuan Zhang, Joya Chen, Shraman Pramanick, Difei Gao, Alex Jinpeng Wang, Rui Yan, and Mike Zheng Shou · 2023
Closest in time.
Multi-modal representation learning with text-driven soft masks
Jaeyoo Park and Bohyung Han · 2023
Closest in time.
An outlook into the future of egocentric vision
Chiara Plizzari, Gabriele Goletto, Antonino Furnari, Siddhant Bansal, Francesco Ragusa, Giovanni Maria Farinella, Dima Damen, and Tatiana Tommasi · 2023
Closest in time.
Steps: Self-supervised key step extraction from unlabeled procedural videos
Anshul Shah, Benjamin Lundell, Harpreet Sawhney, and Rama Chellappa · 2023
Closest in time.
Perceiver-vl: Efficient vision-and-language modeling with iterative latent attention
Zineng Tang, Jaemin Cho, Jie Lei, and Mohit Bansal · 2023
Closest in time.
All in one: Exploring unified video-language pre-training
Jinpeng Wang, Yixiao Ge, Rui Yan, Yuying Ge, Kevin Qinghong Lin, Satoshi Tsutsui, Xudong Lin, Guanyu Cai, Jianping Wu, Ying Shan, et al · 2023
Closest in time.
Position-guided text prompt for vision-language pre-training
Jinpeng Wang, Pan Zhou, Mike Zheng Shou, and Shuicheng Yan · 2023
Closest in time.
Image as a foreign language: Beit pretraining for vision and vision-language tasks
Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, et al · 2023
Closest in time.
Learning video representations from large language models
Yue Zhao, Ishan Misra, Philipp Krähenbühl, and Rohit Girdhar · 2023
Closest in time.