Fetching the paper…
Reading the bibliography…
This paper introduces ChinaOpen, a dataset sourced from Bilibili, a popular Chinese video-sharing website, for open-world multimodal learning.
A Short Note on the Kinetics-700 Human Action Dataset
Joao Carreira, Eric Noland, Chloe Hillier, and Andrew Zisserman. 2022 · 1907
Earlier work this paper cites.
Learning Social Tag Relevance by Neighbor Voting
Xirong Li, Cees G. M. Snoek, and Marcel Worring. 2009 · 2009
Earlier work this paper cites.
Collecting Highly Parallel Data for Paraphrase Evaluation. In CVPR
David Chen and William Dolan. 2011 · 2011
Earlier work this paper cites.
HMDB: A large video database for human motion recognition. In ICCV
Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre. 2011 · 2011
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. 2012 · 2012
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks. In CVPR
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context. In ECCV
Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár. 2014 · 2014
Earlier work this paper cites.
TGIF: A New Dataset and Benchmark on Animated Gif Description. In CVPR
Yuncheng Li, Yale Song, Liangliang Cao, Joel R. Tetreault, Larry Goldberg, Alejandro Jaimes, and Jiebo Luo. 2015 · 2015
Earlier work this paper cites.
Socializing the Semantic Gap: A Comparative Survey on Image Tag Assignment, Refinement, and Retrieval
Xirong Li, Tiberio Uricchio, Lamberto Ballan, Marco Bertini, Cees G. M. Snoek, and Alberto Del Bimbo. 2016 · 2016
Earlier work this paper cites.
MSR-VTT: A Large Video Description Dataset for Bridging Video And Language. In CVPR
Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016 · 2016
Earlier work this paper cites.
The Kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Earlier work this paper cites.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei. 2017b · 2017
Earlier work this paper cites.
Places: A 10 million Image Database for Scene Recognition
Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba. 2017 · 2017
Cited alongside, same era.
A Short Note about Kinetics-600
Joao Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman. 2018 · 2018
Cited alongside, same era.
LVIS: A Dataset for Large Vocabulary Instance Segmentation. In CVPR
Agrim Gupta, Piotr Dollar, and Ross Girshick. 2019 · 2019
Cited alongside, same era.
COCO-CN for Cross-Lingual Image Tagging, Captioning and Retrieval
Xirong Li, Chaoxi Xu, Xiaoxu Wang, Weiyu Lan, Zhengxiong Jia, Gang Yang, and Jieping Xu. 2019 · 2019
Cited alongside, same era.
HowTo100M: Learning a text-video embedding by watching hundred Million Narrated video clips. In ICCV
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. 2019 · 2019
The Stem Cell Hypothesis: Dilemma behind Multi-Task Learning with Transformer Encoders. In EMNLP
Han He and Jinho D. Choi. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision. In ICML
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Flamingo: A visual language model for few-shot learning. In NeurIPS
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Later among the works it cites.
Lightweight Attentional Feature Fusion: A New Baseline for Text-to-Video Retrieval. In ECCV
Fan Hu, Aozhu Chen, Ziyue Wang, Fangming Zhou, Jianfeng Dong, and Xirong Li. 2022 · 2022
Later among the works it cites.
Video Swin Transformer. In CVPR
Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Annotating objects and relations in user-generated videos. In ICMR
Xindi Shang, Donglin Di, Junbin Xiao, Yu Cao, Xun Yang, and Tat-Seng Chua. 2019 · 2019
Cited alongside, same era.
Objects365: A Large-Scale, High-Quality Dataset for Object Detection. In ICCV
Shuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng, Gang Yu, Xiangyu Zhang, Jing Li, and Jian Sun. 2019 · 2019
Cited alongside, same era.
VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research. In ICCV
Xin Wang, Jiawei Wu, Junkun Chen, Lei Li, Yuan-Fang Wang, and William Yang Wang. 2019 · 2019
Cited alongside, same era.
Large scale holistic video understanding. In ECCV
Ali Diba, Mohsen Fayyaz, Vivek Sharma, Manohar Paluri, Jürgen Gall, Rainer Stiefelhagen, and Luc Van Gool. 2020 · 2020
Cited alongside, same era.
The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale
Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, Tom Duerig, and Vittorio Ferrari. 2020 · 2020
Cited alongside, same era.
Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval. In ICCV
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. 2021 · 2021
Cited alongside, same era.
Multi-Level Visual Representation with Semantic-Reinforced Learning for Video Captioning. In ACMMM
Chengbo Dong, Xinru Chen, Aozhu Chen, Fan Hu, Zihan Wang, and Xirong Li. 2021 · 2021
Cited alongside, same era.
CLIP4Clip: An empirical study of CLIP for end to end video clip retrieval and captioning
Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li. 2022 · 2022
Later among the works it cites.
X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text Retrieval. In ACMMM
Yiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan, Ji Zhang, and Rongrong Ji. 2022 · 2022
Later among the works it cites.
Auto-captions on GIF: A Large-scale Video-sentence Dataset for Vision-language Pre-training. In ACMMM
Yingwei Pan, Yehao Li, Jianjie Luo, Jun Xu, Ting Yao, and Tao Mei. 2022 · 2022
Later among the works it cites.
ERNIE-ViL 2.0: Multi-view Contrastive Learning for Image-Text Pre-training
Bin Shan, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. 2022 · 2022
Later among the works it cites.
Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese
An Yang, Junshu Pan, Junyang Lin, Rui Men, Yichang Zhang, Jingren Zhou, and Chang Zhou. 2022 · 2022
Later among the works it cites.
Fengshenbang 1.0: Being the Foundation of Chinese Cognitive Intelligence
Jiaxing Zhang, Ruyi Gan, Junjie Wang, Yuxiang Zhang, Lin Zhang, Ping Yang, Xinyu Gao, Ziwei Wu, Xiaoqun Dong, Junqing He, Jianheng Zhuo, Qi Yang, Yongfeng Huang, Xiayu Li, Yanghan Wu, Junyu Lu, Xinyu Zhu, Weifeng Chen, Ting Han, Kunhao Pan, Rui Wang, Hao Wang, Xiaojun Wu, Zhongshen Zeng, and Chongpei Chen. 2022 · 2022
Later among the works it cites.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023 · 2023
Closest in time.