Fetching the paper…
Reading the bibliography…
This paper presents OmniVL, a new foundation model to support both image-language and video-language tasks using one universal architecture.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
V. Ordonez, G. Kulkarni, and T. Berg · 2011
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
J. Xu, T. Mei, T. Yao, and Y. Rui · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
L. Anne Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell · 2017
Earlier work this paper cites.
The" something something" video database for learning and evaluating visual common sense
R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag, et al · 2017
Earlier work this paper cites.
The kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
D. Xu, Z. Zhao, J. Xiao, F. Wu, H. Zhang, X. He, and Y. Zhuang · 2017
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
P. Sharma, N. Ding, S. Goodman, and R. Soricut · 2018
Earlier work this paper cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
S. Xie, C. Sun, J. Huang, Z. Tu, and K. Murphy · 2018
Earlier work this paper cites.
Towards automatic learning of procedures from web instructional videos
L. Zhou, C. Xu, and J. J. Corso · 2018
Earlier work this paper cites.
End-to-end dense video captioning with masked transformer
L. Zhou, Y. Zhou, J. J. Corso, R. Socher, and C. Xiong · 2018
Earlier work this paper cites.
Nocaps: Novel object captioning at scale
H. Agrawal, K. Desai, Y. Wang, X. Chen, R. Jain, M. Johnson, D. Batra, D. Parikh, S. Lee, and P. Anderson · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Rethinking imagenet pre-training
K. He, R. Girshick, and P. Dollár · 2019
Earlier work this paper cites.
A case study on combining asr and visual features for generating instructional video captions
J. Hessel, B. Pang, Z. Zhu, and R. Soricut · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
J. Lu, D. Batra, D. Parikh, and S. Lee · 2019
Earlier work this paper cites.
Videobert: A joint model for video and language representation learning
C. Sun, A. Myers, C. Vondrick, K. Murphy, and C. Schmid · 2019
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
H. Tan and M. Bansal · 2019
Earlier work this paper cites.
Self-supervised multimodal versatile networks
J.-B. Alayrac, A. Recasens, R. Schneider, R. Arandjelović, J. Ramapuram, J. De Fauw, L. Smaira, S. Dieleman, and A. Zisserman · 2020
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton · 2020
Earlier work this paper cites.
Uniter: Universal image-text representation learning
Y.-C. Chen, L. Li, L. Yu, A. E. Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu · 2020
Earlier work this paper cites.
Randaugment: Practical automated data augmentation with a reduced search space
E. D. Cubuk, B. Zoph, J. Shlens, and Q. V. Le · 2020
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick · 2020
Cited alongside, same era.
Big transfer (bit): General visual representation learning
A. Kolesnikov, L. Beyer, X. Zhai, J. Puigcerver, J. Yung, S. Gelly, and N. Houlsby · 2020
Cited alongside, same era.
Hero: Hierarchical encoder for video+ language omni-representation pre-training
L. Li, Y.-C. Chen, Y. Cheng, Z. Gan, L. Yu, and J. Liu · 2020
Cited alongside, same era.
Unimo: Towards unified-modal understanding and generation via cross-modal contrastive learning
W. Li, C. Gao, G. Niu, X. Xiao, H. Liu, J. Liu, H. Wu, and H. Wang · 2020
Cited alongside, same era.
Oscar: Object-semantics aligned pre-training for vision-language tasks
X. Li, X. Yin, C. Li, P. Zhang, X. Hu, L. Zhang, L. Wang, H. Hu, L. Dong, F. Wei, et al · 2020
Cited alongside, same era.
Improve unsupervised pretraining for few-label transfer
S. Li, D. Chen, Y. Chen, L. Yuan, L. Zhang, Q. Chu, B. Liu, and N. Yu · 2021
Later among the works it cites.
Pretrained transformers as universal computation engines
K. Lu, A. Grover, P. Abbeel, and I. Mordatch · 2021
Later among the works it cites.
Thinking fast and slow: Efficient text-to-visual retrieval with transformers
A. Miech, J.-B. Alayrac, I. Laptev, J. Sivic, and A. Zisserman · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Later among the works it cites.
Actionclip: A new paradigm for video action recognition
M. Wang, J. Xing, and Y. Liu · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Luo, L. Ji, B. Shi, H. Huang, N. Duan, T. Li, J. Li, T. Bharti, and M. Zhou · 2020
Cited alongside, same era.
End-to-end learning of visual representations from uncurated instructional videos
A. Miech, J.-B. Alayrac, L. Smaira, I. Laptev, J. Sivic, and A. Zisserman · 2020
Cited alongside, same era.
Support-set bottlenecks for video-text representation learning
M. Patrick, P.-Y. Huang, Y. Asano, F. Metze, A. Hauptmann, J. Henriques, and A. Vedaldi · 2020
Cited alongside, same era.
Vl-bert: Pre-training of generic visual-linguistic representations
W. Su, X. Zhu, Y. Cao, B. Li, L. Lu, F. Wei, and J. Dai · 2020
Cited alongside, same era.
Unified vision-language pre-training for image captioning and vqa
L. Zhou, H. Palangi, L. Zhang, H. Hu, J. Corso, and J. Gao · 2020
Cited alongside, same era.
Actbert: Learning global-local video-text representations
L. Zhu and Y. Yang · 2020
Cited alongside, same era.
A comprehensive study of deep video action recognition
Y. Zhu, X. Li, C. Liu, M. Zolfaghari, Y. Xiong, C. Wu, Z. Zhang, J. Tighe, R. Manmatha, and M. Li · 2020
Cited alongside, same era.
Vlmo: Unified vision-language pre-training with mixture-of-modality-experts
W. Wang, H. Bao, L. Dong, and F. Wei · 2021
Later among the works it cites.
Videoclip: Contrastive pre-training for zero-shot video-text understanding
H. Xu, G. Ghosh, P.-Y. Huang, D. Okhonko, A. Aghajanyan, F. Metze, L. Zettlemoyer, and C. Feichtenhofer · 2021
Later among the works it cites.
Just ask: Learning to answer questions from millions of narrated videos
A. Yang, A. Miech, J. Sivic, I. Laptev, and C. Schmid · 2021
Later among the works it cites.
Florence: A new foundation model for computer vision
L. Yuan, D. Chen, Y.-L. Chen, N. Codella, X. Dai, J. Gao, H. Hu, X. Huang, B. Li, C. Li, et al · 2021
Later among the works it cites.
Merlot: Multimodal neural script knowledge models
R. Zellers, X. Lu, J. Hessel, Y. Yu, J. S. Park, J. Cao, A. Farhadi, and Y. Choi · 2021
Later among the works it cites.
Co-training transformer with videos and images improves action recognition
B. Zhang, J. Yu, C. Fifty, W. Han, A. M. Dai, R. Pang, and F. Sha · 2021
Later among the works it cites.
Vinvl: Revisiting visual representations in vision-language models
P. Zhang, X. Li, X. Hu, J. Yang, L. Zhang, L. Wang, Y. Choi, and J. Gao · 2021
Later among the works it cites.
X. Zhu, J. Zhu, H. Li, X. Wu, X. Wang, H. Li, X. Wang, and J. Dai · 2021
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Closest in time.
Beit: Bert pre-training of image transformers
H. Bao, L. Dong, and F. Wei · 2022
Closest in time.
Cross modal retrieval with querybank normalisation
S.-V. Bogolin, I. Croitoru, H. Jin, Y. Liu, and S. Albanie · 2022
Closest in time.
Cswin transformer: A general vision transformer backbone with cross-shaped windows
X. Dong, J. Bao, D. Chen, W. Zhang, N. Yu, L. Yuan, D. Chen, and B. Guo · 2022
Closest in time.
Bootstrapped masked autoencoders for vision bert pretraining
X. Dong, J. Bao, T. Zhang, D. Chen, W. Zhang, L. Yuan, D. Chen, F. Wen, and N. Yu · 2022
Closest in time.
Align and prompt: Video-and-language pre-training with entity prompts
D. Li, J. Li, H. Li, J. C. Niebles, and S. C. Hoi · 2022
Closest in time.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
J. Li, D. Li, C. Xiong, and S. Hoi · 2022
Closest in time.
Flava: A foundational language and vision alignment model
A. Singh, R. Hu, V. Goswami, G. Couairon, W. Galuba, M. Rohrbach, and D. Kiela · 2022
Closest in time.
Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
P. Wang, A. Yang, R. Men, J. Lin, S. Bai, Z. Li, J. Ma, C. Zhou, J. Zhou, and H. Yang · 2022
Closest in time.
Bevt: Bert pretraining of video transformers
R. Wang, D. Chen, Z. Wu, Y. Chen, X. Dai, M. Liu, Y.-G. Jiang, L. Zhou, and L. Yuan · 2022
Closest in time.
SimVLM: Simple visual language model pretraining with weak supervision
Z. Wang, J. Yu, A. W. Yu, Z. Dai, Y. Tsvetkov, and Y. Cao · 2022
Closest in time.
Unified contrastive learning in image-text-label space
J. Yang, C. Li, P. Zhang, B. Xiao, C. Liu, L. Yuan, and J. Gao · 2022
Closest in time.
Coca: Contrastive captioners are image-text foundation models
J. Yu, Z. Wang, V. Vasudevan, L. Yeung, M. Seyedhosseini, and Y. Wu · 2022
Closest in time.