Fetching the paper…
Reading the bibliography…
Recent years have witnessed a big convergence of language, vision, and multi-modal pretraining.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
Chen, D. and Dolan, W. B · 2011
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Ordonez, V., Kulkarni, G., and Berg, T. L · 2011
Earlier work this paper cites.
Unimo: Towards unified-modal understanding and generation via cross-modal contrastive learning
Li, W., Gao, C., Niu, G., Xiao, X., Liu, H., Liu, J., Wu, H., and Wang, H · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A. and Fei-Fei, L · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J., and Lazebnik, S · 2015
Earlier work this paper cites.
A dataset for movie description
Rohrbach, A., Rohrbach, M., Tandon, N., and Schiele, B · 2015
Earlier work this paper cites.
A neural attention model for abstractive sentence summarization
Rush, A. M., Chopra, S., and Weston, J · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R. S., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Mao, J., Huang, J., Toshev, A., Camburu, O., Yuille, A. L., and Murphy, K · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Xu, J., Mei, T., Yao, T., and Rui, Y · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Yu, L., Poirson, P., Yang, S., Berg, A. C., and Berg, T. L · 2016
Earlier work this paper cites.
Vqa: Visual question answering
Agrawal, A., Lu, J., Antol, S., Mitchell, M., Zitnick, C. L., Parikh, D., and Batra, D · 2017
Earlier work this paper cites.
Localizing moments in video with natural language
Anne Hendricks, L., Wang, O., Shechtman, E., Sivic, J., Darrell, T., and Russell, B · 2017
Earlier work this paper cites.
Mask r-cnn
He, K., Gkioxari, G., Dollár, P., and Girshick, R. B · 2017
Earlier work this paper cites.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Jang, Y., Song, Y., Yu, Y., Kim, Y., and Kim, G · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, A., Suleyman, M., and Zisserman, A · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.-J., Shamma, D. A., et al · 2017
Earlier work this paper cites.
Inception-v4, inception-resnet and the impact of residual connections on learning
Szegedy, C., Ioffe, S., Vanhoucke, V., and Alemi, A. A · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Xu, D., Zhao, Z., Xiao, J., Wu, F., Zhang, H., He, X., and Zhuang, Y · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Earlier work this paper cites.
Unified language model pre-training for natural language understanding and generation
Dong, L., Yang, N., Wang, W., Wei, F., Liu, X., Wang, Y., Gao, J., Zhou, M., and Hon, H · 2019
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Li, L. H., Yatskar, M., Yin, D., Hsieh, C.-J., and Chang, K.-W · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J., Batra, D., Parikh, D., and Lee, S · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Köpf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Cited alongside, same era.
Objects365: A large-scale, high-quality dataset for object detection
Shao, S., Li, Z., Zhang, T., Peng, C., Yu, G., Zhang, X., Li, J., and Sun, J · 2019
Cited alongside, same era.
MASS: masked sequence to sequence pre-training for language generation
Song, K., Tan, X., Qin, T., Lu, J., and Liu, T · 2019
Cited alongside, same era.
Vl-bert: Pre-training of generic visual-linguistic representations
Su, W., Zhu, X., Cao, Y., Li, B., Lu, L., Wei, F., and Dai, J · 2019
Cited alongside, same era.
Lxmert: Learning cross-modality encoder representations from transformers
Tan, H. and Bansal, M · 2019
Movinets: Mobile video networks for efficient video recognition
Kondratyuk, D., Yuan, L., Li, Y., Zhang, L., Tan, M., Brown, M., and Gong, B · 2021
Later among the works it cites.
Cbnet: A composite backbone network architecture for object detection
Liang, T., Chu, X., Liu, Y., Wang, Y., Tang, Z., Chu, W., Chen, J., and Ling, H · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Later among the works it cites.
Tokenlearner: What can 8 learned tokens do for images and videos?
Ryoo, M. S., Piergiovanni, A. J., Arnab, A., Dehghani, M., and Angelova, A · 2021
Later among the works it cites.
Flava: A foundational language and vision alignment model
Singh, A., Hu, R., Goswami, V., Couairon, G., Galuba, W., Rohrbach, M., and Kiela, D · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
Tan, M. and Le, Q. V · 2019
Cited alongside, same era.
Xlnet: Generalized autoregressive pretraining for language understanding
Yang, Z., Dai, Z., Yang, Y., Carbonell, J. G., Salakhutdinov, R., and Le, Q. V · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
End-to-end object detection with transformers
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S · 2020
Cited alongside, same era.
Uniter: Universal image-text representation learning
Chen, Y.-C., Li, L., Yu, L., El Kholy, A., Ahmed, F., Gan, Z., Cheng, Y., and Liu, J · 2020
Cited alongside, same era.
ELECTRA: pre-training text encoders as discriminators rather than generators
Clark, K., Luong, M., Le, Q. V., and Manning, C. D · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Cited alongside, same era.
Clip4caption: Clip for video caption
Tang, M., Wang, Z., Liu, Z., Rao, F., Li, D., and Li, X · 2021
Later among the works it cites.
E2E-VLP: end-to-end vision-language pre-training enhanced by visual learning
Xu, H., Yan, M., Li, C., Bi, B., Huang, S., Xiao, W., and Huang, F · 2021
Later among the works it cites.
Florence: A new foundation model for computer vision
Yuan, L., Chen, D., Chen, Y.-L., Codella, N., Dai, X., Gao, J., Hu, H., Huang, X., Li, B., Li, C., et al · 2021
Later among the works it cites.
Merlot: Multimodal neural script knowledge models
Zellers, R., Lu, X., Hessel, J., Yu, Y., Park, J. S., Cao, J., Farhadi, A., and Choi, Y · 2021
Later among the works it cites.
Scaling vision transformers
Zhai, X., Kolesnikov, A., Houlsby, N., and Beyer, L · 2021
Later among the works it cites.
Zhu, X., Zhu, J., Li, H., Wu, X., Wang, X., Li, H., Wang, X., and Dai, J · 2021
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., Ring, R., Rutherford, E., Cabi, S., Han, T., Gong, Z., Samangooei, S., Monteiro, M., Menick, J., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Binkowski, M., Barreira, R., Vinyals, O., Zisserman, A., and Simonyan, K · 2022
Later among the works it cites.
Bridging video-text retrieval with multiple choice questions
Ge, Y., Ge, Y., Liu, X., Li, D., Shan, Y., Qie, X., and Luo, P · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Doll’ar, P., and Girshick, R. B · 2022
Later among the works it cites.
Revealing single frame bias for video-and-language learning
Lei, J., Berg, T. L., and Bansal, M · 2022
Later among the works it cites.
Grounded language-image pre-training
Li, L. H., Zhang, P., Zhang, H., Yang, J., Li, C., Zhong, Y., Wang, L., Yuan, L., Zhang, L., Hwang, J.-N., Chang, K.-W., and Gao, J · 2022
Later among the works it cites.
Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning
Liang, W., Zhang, Y., Kwon, Y., Yeung, S., and Zou, J · 2022
Later among the works it cites.
Swinbert: End-to-end transformers with sparse attention for video captioning
Lin, K., Li, L., Lin, C.-C., Ahmed, F., Gan, Z., Liu, Z., Lu, Y., and Wang, L · 2022
Later among the works it cites.
Swin transformer v2: Scaling up capacity and resolution
Liu, Z., Hu, H., Lin, Y., Yao, Z., Xie, Z., Wei, Y., Ning, J., Cao, Y., Zhang, Z., Dong, L., Wei, F., and Guo, B · 2022
Later among the works it cites.
A convnet for the 2020s
Liu, Z., Mao, H., Wu, C., Feichtenhofer, C., Darrell, T., and Xie, S · 2022
Later among the works it cites.
Video swin transformer
Liu, Z., Ning, J., Cao, Y., Wei, Y., Zhang, Z., Lin, S., and Hu, H · 2022
Later among the works it cites.
Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning
Luo, H., Ji, L., Zhong, M., Chen, Y., Lei, W., Duan, N., and Li, T · 2022
Later among the works it cites.
X-clip: End-to-end multi-grained contrastive learning for video-text retrieval
Ma, Y., Xu, G., Sun, X., Yan, M., Zhang, J., and Ji, R · 2022
Later among the works it cites.
End-to-end generative pretraining for multimodal video captioning
Seo, P. H., Nagrani, A., Arnab, A., and Schmid, C · 2022
Later among the works it cites.
Deit iii: Revenge of the vit
Touvron, H., Cord, M., and J’egou, H · 2022
Later among the works it cites.
Groupvit: Semantic segmentation emerges from text supervision
Xu, J., Mello, S. D., Liu, S., Byeon, W., Breuel, T., Kautz, J., and Wang, X · 2022
Later among the works it cites.
Video-text modeling with zero-shot transfer from contrastive captioners
Yan, S., Zhu, T., Wang, Z., Cao, Y., Zhang, M., Ghosh, S., Wu, Y., and Yu, J · 2022
Later among the works it cites.
Zero-shot video question answering via frozen bidirectional language models
Yang, A., Miech, A., Sivic, J., Laptev, I., and Schmid, C · 2022
Later among the works it cites.
Hitea: Hierarchical temporal-aware video-language pre-training
Ye, Q., Xu, G., Yan, M., Xu, H., Qian, Q., Zhang, J. C., and Huang, F · 2022
Later among the works it cites.
Coca: Contrastive captioners are image-text foundation models
Yu, J., Wang, Z., Vasudevan, V., Yeung, L., Seyedhosseini, M., and Wu, Y · 2022
Later among the works it cites.