Fetching the paper…
Reading the bibliography…
We study joint video and language (VL) pre-training to enable cross-modality learning and benefit plentiful downstream VL tasks.
Personalized video recommendation through tripartite graph propagation
Bisheng Chen, Jingdong Wang, Qinghua Huang, and Tao Mei · 2012
Earlier work this paper cites.
Image tag refinement with view-dependent concept representations
Jianlong Fu, Jinqiao Wang, Yong Rui, Xin-Jing Wang, Tao Mei, and Hanqing Lu · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Tagging personal photos with transfer deep learning
Jianlong Fu, Tao Mei, Kuiyuan Yang, Hanqing Lu, and Yong Rui · 2015
Earlier work this paper cites.
Deep learning face attributes in the wild
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Translating videos to natural language using deep recurrent neural networks
Subhashini Venugopalan, Huijuan Xu, Jeff Donahue, Marcus Rohrbach, Raymond J Mooney, and Kate Saenko · 2015
Earlier work this paper cites.
Youtube-8m: A large-scale video classification benchmark
Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Hadamard product for low-rank bilinear pooling
Jin-Hwa Kim, Kyoung-Woon On, Woosang Lim, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang · 2016
Earlier work this paper cites.
Movie description
Anna Rohrbach, Atousa Torabi, Marcus Rohrbach, Niket Tandon, Christopher Joseph Pal, H. Larochelle, Aaron C. Courville, and Bernt Schiele · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Advances in deep learning approaches for image tagging
Jianlong Fu and Yong Rui · 2017
Earlier work this paper cites.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang · 2017
Earlier work this paper cites.
End-to-end concept word detection for video captioning, retrieval, and question answering
Youngjae Yu, Hyungjin Ko, Jongwook Choi, and Gunhee Kim · 2017
Earlier work this paper cites.
Motion-appearance co-memory networks for video question answering
Jiyang Gao, Runzhou Ge, Kan Chen, and Ram Nevatia · 2018
Earlier work this paper cites.
Progressive growing of gans for improved quality, stability, and variation
Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen · 2018
Earlier work this paper cites.
The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english
Steven R Livingstone and Frank A Russo · 2018
Earlier work this paper cites.
Learning a text-video embedding from incomplete and heterogeneous data
Antoine Miech, Ivan Laptev, and Josef Sivic · 2018
Earlier work this paper cites.
How2: A large-scale dataset for multimodal language understanding
Ramon Sanabria, Ozan Caglayan, Shruti Palaskar, Desmond Elliott, Loïc Barrault, Lucia Specia, and Florian Metze · 2018
Earlier work this paper cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy · 2018
Cited alongside, same era.
A joint sequence fusion model for video question answering and retrieval
Youngjae Yu, Jongseok Kim, and Gunhee Kim · 2018
Cited alongside, same era.
Cross-modal and hierarchical modeling of video and text
Bowen Zhang, Hexiang Hu, and Fei Sha · 2018
Cited alongside, same era.
S3d: Single shot multi-span detector via fully 3d convolutional network
Da Zhang, Xiyang Dai, Xin Wang, and Yuan-Fang Wang · 2018
Cited alongside, same era.
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang · 2018
Cited alongside, same era.
Towards automatic learning of procedures from web instructional videos
End-to-end learning of visual representations from uncurated instructional videos
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman · 2020
Later among the works it cites.
Conditional image generation and manipulation for user-specified content
David Stap, Maurits Bleeker, Sarah Ibrahimi, and Maartje ter Hoeve · 2020
Later among the works it cites.
Learning semantic-aware normalization for generative adversarial networks
Heliang Zheng, Jianlong Fu, Yanhong Zeng, Jiebo Luo, and Zheng-Jun Zha · 2020
Later among the works it cites.
Actbert: Learning global-local video-text representations
Linchao Zhu and Yi Yang · 2020
Later among the works it cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman · 2021
Closest in time.
Is space-time attention all you need for video understanding?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Luowei Zhou, Chenliang Xu, and Jason J. Corso · 2018
Cited alongside, same era.
Neural storyboard artist: Visualizing stories with coherent image sequences
Shizhe Chen, Bei Liu, Jianlong Fu, Ruihua Song, Qin Jin, Pingping Lin, Xiaoyu Qi, Chunting Wang, and Jin Zhou · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Heterogeneous memory enhanced multimodal attention model for video question answering
Chenyou Fan, Xiaofan Zhang, Shu Zhang, Wensheng Wang, Chi Zhang, and Heng Huang · 2019
Cited alongside, same era.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Cited alongside, same era.
A style-based generator architecture for generative adversarial networks
Tero Karras, Samuli Laine, and Timo Aila · 2019
Cited alongside, same era.
Emotion reinforced visual storytelling
Nanxing Li, Bei Liu, Zhizhong Han, Yu-Shen Liu, and Jianlong Fu · 2019
Cited alongside, same era.
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Closest in time.
Datasheets for datasets
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford · 2021
Closest in time.
Unifying multimodal transformer for bi-directional image and text generation
Yupan Huang, Hongwei Xue, Bei Liu, and Yutong Lu · 2021
Closest in time.
Seeing out of the box: End-to-end pre-training for vision-language representation learning
Zhicheng Huang, Zhaoyang Zeng, Yupan Huang, Bei Liu, Dongmei Fu, and Jianlong Fu · 2021
Closest in time.
Less is more: Clipbert for video-and-language learning via sparse sampling
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu · 2021
Closest in time.
Align before fuse: Vision and language representation learning with momentum distillation
Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Deepak Gotmare, Shafiq R. Joty, Caiming Xiong, and Steven C. H. Hoi · 2021
Closest in time.
Styleclip: Text-driven manipulation of stylegan imagery
Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, and Dani Lischinski · 2021
Closest in time.
Support-set bottlenecks for video-text representation learning
Mandela Patrick, Po-Yao Huang, Yuki Asano, Florian Metze, Alexander G Hauptmann, Joao F. Henriques, and Andrea Vedaldi · 2021
Closest in time.
Encoding in style: a stylegan encoder for image-to-image translation
Elad Richardson, Yuval Alaluf, Or Patashnik, Yotam Nitzan, Yaniv Azar, Stav Shapiro, and Daniel Cohen-Or · 2021
Closest in time.
Image super-resolution via iterative refinement
Chitwan Saharia, Jonathan Ho, William Chan, Tim Salimans, David J Fleet, and Mohammad Norouzi · 2021
Closest in time.
Decembert: Learning from noisy instructional videos via dense captions and entropy minimization
Zineng Tang, Jie Lei, and Mohit Bansal · 2021
Closest in time.
Designing an encoder for stylegan image manipulation
Omer Tov, Yuval Alaluf, Yotam Nitzan, Or Patashnik, and Daniel Cohen-Or · 2021
Closest in time.
Tedigan: Text-guided diverse face image generation and manipulation
Weihao Xia, Yujiu Yang, Jing-Hao Xue, and Baoyuan Wu · 2021
Closest in time.
Vlm: Task-agnostic video-language model pre-training for video understanding
Hu Xu, Gargi Ghosh, Po-Yao Huang, Prahal Arora, Masoumeh Aminzadeh, Christoph Feichtenhofer, Florian Metze, and Luke Zettlemoyer · 2021
Closest in time.
Videoclip: Contrastive pre-training for zero-shot video-text understanding
Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer · 2021
Closest in time.
Probing inter-modality: Visual parsing with self-attention for vision-and-language pre-training
Hongwei Xue, Yupan Huang, Bei Liu, Houwen Peng, Jianlong Fu, Houqiang Li, and Jiebo Luo · 2021
Closest in time.
Merlot: Multimodal neural script knowledge models
Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi · 2021
Closest in time.
Improving visual quality of image synthesis by a token-based generator with transformers
Yanhong Zeng, Huan Yang, Hongyang Chao, Jianbo Wang, and Jianlong Fu · 2021
Closest in time.