Fetching the paper…
Reading the bibliography…
Video-Language Pretraining (VLP), which aims to learn transferable representation to advance a wide range of video-text downstream tasks, has recently received increasing attention.
A user attention model for video summarization
Yu-Fei Ma, Lie Lu, Hong-Jiang Zhang, and Mingjing Li · 2002
Earlier work this paper cites.
Bidirectional lstm networks for improved phoneme classification and recognition
Alex Graves, Santiago Fernández, and Jürgen Schmidhuber · 2005
Earlier work this paper cites.
Discovering important people and objects for egocentric video summarization
Yong Jae Lee, Joydeep Ghosh, and Kristen Grauman · 2012
Earlier work this paper cites.
Social interactions: A first-person perspective
Alircza Fathi, Jessica K Hodgins, and James M Rehg · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Delving into egocentric actions
Yin Li, Zhefan Ye, and James M Rehg · 2015
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Earlier work this paper cites.
Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning
Long Chen, Hanwang Zhang, Jun Xiao, Liqiang Nie, Jian Shao, Wei Liu, and Tat-Seng Chua · 2017
Earlier work this paper cites.
TALL: Temporal Activity Localization via Language Query
Gao Jiyang, Sun Chen, Yang Zhenheng, Nevatia, Ram · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
A joint sequence fusion model for video question answering and retrieval
Youngjae Yu, Jongseok Kim, and Gunhee Kim · 2018
Earlier work this paper cites.
Reconstruction network for video captioning
Bairui Wang, Lin Ma, Wei Zhang, and Wei Liu · 2018
Earlier work this paper cites.
End-to-end dense video captioning with masked transformer
Luowei Zhou, Yingbo Zhou, Jason J Corso, Richard Socher, and Caiming Xiong · 2018
Earlier work this paper cites.
When will you do what?-anticipating temporal occurrences of activities
Yazan Abu Farha, Alexander Richard, and Juergen Gall · 2018
Earlier work this paper cites.
Towards automatic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason J Corso · 2018
Cited alongside, same era.
Charades-ego: A large-scale dataset of paired third and first person videos
Gunnar A Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Aaron Van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Cited alongside, same era.
Localizing moments in video with temporal language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2018
Cited alongside, same era.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy · 2018
Cited alongside, same era.
Span-based localizing network for natural language video localization
Hao Zhang, Aixin Sun, Wei Jing, and Joey Tianyi Zhou · 2020
Later among the works it cites.
Unsupervised and semi-supervised domain adaptation for action recognition from drones
Jinwoo Choi, Gaurav Sharma, Manmohan Chandraker, and Jia-Bin Huang · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Later among the works it cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman · 2021
Later among the works it cites.
Less is more: Clipbert for video-and-language learning via sparse sampling
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Actor and observer: Joint modeling of first and third-person videos
Gunnar A Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari · 2018
Cited alongside, same era.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Cited alongside, same era.
Epic-fusion: Audio-visual temporal binding for egocentric action recognition
Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen · 2019
Cited alongside, same era.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid · 2019
Cited alongside, same era.
Temporal localization of moments in video collections with natural language
Victor Escorcia, Mattia Soldan, Josef Sivic, Bernard Ghanem, and Bryan Russell · 2019
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Cited alongside, same era.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Cited alongside, same era.
Thinking fast and slow: Efficient text-to-visual retrieval with transformers
Antoine Miech, Jean-Baptiste Alayrac, Ivan Laptev, Josef Sivic, and Andrew Zisserman · 2021
Later among the works it cites.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Later among the works it cites.
On semantic similarity in video retrieval
Michael Wray, Hazel Doughty, and Dima Damen · 2021
Later among the works it cites.
Vlg-net: Video-language graph matching network for video grounding
Mattia Soldan, Mengmeng Xu, Sisi Qu, Jesper Tegner, and Bernard Ghanem · 2021
Later among the works it cites.
Ego-exo: Transferring visual representations from third-person to first-person videos
Yanghao Li, Tushar Nagarajan, Bo Xiong, and Kristen Grauman · 2021
Later among the works it cites.
Video self-stitching graph network for temporal action localization
Chen Zhao, Ali K Thabet, and Bernard Ghanem · 2021
Later among the works it cites.
Detecting moments and highlights in videos via natural language queries
Jie Lei, Tamara L Berg, and Mohit Bansal · 2021
Later among the works it cites.
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al · 2022
Closest in time.
Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al · 2022
Closest in time.
Object-aware video-language pre-training for retrieval
Jinpeng Wang, Yixiao Ge, Guanyu Cai, Rui Yan, Xudong Lin, Ying Shan, Xiaohu Qie, and Mike Zheng Shou · 2022
Closest in time.
Assistq: Affordance-centric question-driven task completion for egocentric assistant
Benita Wong, Joya Chen, You Wu, Stan Weixian Lei, Dongxing Mao, Difei Gao, and Mike Zheng Shou · 2022
Closest in time.
Video swin transformer
Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu · 2022
Closest in time.