Fetching the paper…
Reading the bibliography…
With the recent success of the pre-training technique for NLP and image-linguistic tasks, some video-linguistic pre-training works are gradually developed to improve video-text related downstream tasks.
Unified language model pre-training for natural language understanding and generation
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. 2019 · 1905
Earlier work this paper cites.
Mass: Masked sequence to sequence pre-training for language generation
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2019 · 1905
Earlier work this paper cites.
Contrastive bidirectional transformer for temporal representation learning
Chen Sun, Fabien Baradel, Kevin Murphy, and Cordelia Schmid. 2019a · 1906
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V Le. 2019 · 1906
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training
Gen Li, Nan Duan, Yuejian Fang, Daxin Jiang, and Ming Zhou. 2019a · 1908
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019b · 1908
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal. 2019 · 1908
Earlier work this paper cites.
Unified vision-language pre-training for image captioning and vqa
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason J Corso, and Jianfeng Gao. 2019 · 1909
Earlier work this paper cites.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019 · 1910
Earlier work this paper cites.
Factorized multimodal transformer for multimodal sequential learning
Amir Zadeh, Chengfeng Mao, Kelly Shi, Yiwei Zhang, Paul Pu Liang, Soujanya Poria, and Louis-Philippe Morency. 2019 · 1911
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics
Chin-Yew Lin and Franz Josef Och. 2004 · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Hero: Hierarchical encoder for video+ language omni-representation pre-training
Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu. 2020 · 2005
Earlier work this paper cites.
Video understanding as machine translation
Bruno Korbar, Fabio Petroni, Rohit Girdhar, and Lorenzo Torresani. 2020 · 2006
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S Zemel. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Associating neural word embeddings with deep image representations using fisher vectors
Benjamin Klein, Guy Lev, Gil Sadeh, and Lior Wolf. 2015 · 2015
Earlier work this paper cites.
Deep multi-scale video prediction beyond mean square error
Michael Mathieu, Camille Couprie, and Yann LeCun. 2015 · 2015
Cited alongside, same era.
Unsupervised learning of video representations using lstms
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhudinov. 2015 · 2015
Cited alongside, same era.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Cited alongside, same era.
Unsupervised learning of visual representations using videos
Xiaolong Wang and Abhinav Gupta. 2015 · 2015
Cited alongside, same era.
Unsupervised learning from narrated instruction videos
Jean-Baptiste Alayrac, Piotr Bojanowski, Nishant Agrawal, Josef Sivic, Ivan Laptev, and Simon Lacoste-Julien. 2016 · 2016
Cited alongside, same era.
Deep predictive coding networks for video prediction and unsupervised learning
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
Neuralnetwork-viterbi: A framework for weakly supervised video learning
Alexander Richard, Hilde Kuehne, Ahsan Iqbal, and Juergen Gall. 2018 · 2018
Later among the works it cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy. 2018 · 2018
Later among the works it cites.
A joint sequence fusion model for video question answering and retrieval
Youngjae Yu, Jongseok Kim, and Gunhee Kim. 2018 · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
William Lotter, Gabriel Kreiman, and David Cox. 2016 · 2016
Cited alongside, same era.
Extending long short-term memory for multi-view structured learning
Shyam Sundar Rajagopalan, Louis-Philippe Morency, Tadas Baltrusaitis, and Roland Goecke. 2016 · 2016
Cited alongside, same era.
Learning language-visual embedding for movie understanding with natural-language
Atousa Torabi, Niket Tandon, and Leonid Sigal. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016 · 2016
Cited alongside, same era.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016 · 2016
Cited alongside, same era.
Video captioning and retrieval models with semantic attention
Youngjae Yu, Hyungjin Ko, Jongwook Choi, and Gunhee Kim. 2016 · 2016
Cited alongside, same era.
Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages
Amir Zadeh, Rowan Zellers, Eli Pincus, and Louis-Philippe Morency. 2016 · 2016
Cited alongside, same era.
Tengda Han, Weidi Xie, and Andrew Zisserman. 2019 · 2019
Later among the works it cites.
A case study on combining asr and visual features for generating instructional video captions
Jack Hessel, Bo Pang, Zhenhai Zhu, and Radu Soricut. 2019 · 2019
Later among the works it cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Later among the works it cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. 2019 · 2019
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019 · 2019
Later among the works it cites.
Dense procedure captioning in narrated instructional videos
Botian Shi, Lei Ji, Yaobo Liang, Nan Duan, Peng Chen, Zhendong Niu, and Ming Zhou. 2019 · 2019
Later among the works it cites.
COIN: A large-scale dataset for comprehensive instructional video analysis
Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, and Jie Zhou. 2019 · 2019
Later among the works it cites.
Multimodal transformer for unaligned multimodal language sequences
Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J. Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019 · 2019
Later among the works it cites.
Words can shift: Dynamically adjusting word representations using nonverbal behaviors
Yansen Wang, Ying Shen, Zhun Liu, Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency. 2019 · 2019
Later among the works it cites.
Cross-task weakly supervised learning from instructional videos
Dimitri Zhukov, Jean-Baptiste Alayrac, Ramazan Gokberk Cinbis, David Fouhey, Ivan Laptev, and Josef Sivic. 2019 · 2019
Later among the works it cites.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Closest in time.
End-to-End Learning of Visual Representations from Uncurated Instructional Videos
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman. 2020 · 2020
Closest in time.
VL-BERT: pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020 · 2020
Closest in time.
Actbert: Learning global-local video-text representations
Linchao Zhu and Yi Yang. 2020 · 2020
Closest in time.