Fetching the paper…
Reading the bibliography…
English-based Vision-Language Pre-training (VLP) has achieved great success in various downstream tasks.
Visualizing data using t-sne
Van der Maaten, L. and Hinton, G · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P., Lai, A., Hodosh, M., and Hockenmaier, J · 2014
Earlier work this paper cites.
First language acquisition and classroom language learning: Similarities and differences
Castello, D · 2015
Earlier work this paper cites.
Microsoft COCO captions: Data collection and evaluation server
Chen, X., Fang, H., Lin, T., Vedantam, R., Gupta, S., Dollár, P., and Zitnick, C. L · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A. and Fei-Fei, L · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Multi30k: Multilingual english-german image descriptions
Elliott, D., Frank, S., Sima’an, K., and Specia, L · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Xu, J., Mei, T., Yao, T., and Rui, Y · 2016
Earlier work this paper cites.
Image pivoting for learning multilingual multimodal representations
Gella, S., Sennrich, R., Keller, F., and Lapata, M · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Stair captions: Constructing a large-scale japanese image caption dataset
Yoshikawa, Y., Shigeto, Y., and Takeuchi, A · 2017
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R · 2018
Earlier work this paper cites.
A joint sequence fusion model for video question answering and retrieval
Yu, Y., Kim, J., and Kim, G · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for NLP
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Cited alongside, same era.
Coco-cn for cross-lingual image tagging, captioning, and retrieval
Li, X., Xu, C., Wang, X., Lan, W., Jia, Z., Yang, G., and Xu, J · 2019
Cited alongside, same era.
Do imagenet classifiers generalize to imagenet?
Recht, B., Roelofs, R., Schmidt, L., and Shankar, V · 2019
Cited alongside, same era.
Language-agnostic visual-semantic embeddings
Wehrmann, J., Souza, D. M., Lopes, M. A., and Barros, R. C · 2019
Multilingual clip
Carlsson, F · 2021
Later among the works it cites.
Cross-lingual cross-modal pretraining for multimodal retrieval
Fei, H., Yu, T., and Li, P · 2021
Later among the works it cites.
Multilingual multimodal pre-training for zero-shot cross-lingual transfer of vision-language models
Huang, P.-Y., Patrick, M., Hu, J., Neubig, G., Metze, F., and Hauptmann, A. G · 2021
Later among the works it cites.
Wenlan: Bridging vision and language by large-scale multi-modal pre-training, 2021
Huo, Y., Zhang, M., Liu, G., Lu, H., Gao, Y., et al · 2021
Later among the works it cites.
MURAL: Multimodal, multitask representations across languages
Jain, A., Guo, M., Srinivasan, K., Chen, T., Kudugunta, S., Jia, C., Yang, Y., and Baldridge, J · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Towards zero-shot cross-lingual image retrieval
Aggarwal, P. and Kale, A · 2020
Cited alongside, same era.
On the cross-lingual transferability of monolingual representations
Artetxe, M., Ruder, S., and Yogatama, D · 2020
Cited alongside, same era.
Learning to scale multilingual representations for vision-language tasks
Burns, A., Kim, D., Wijaya, D., Saenko, K., and Plummer, B. A · 2020
Cited alongside, same era.
Uniter: Universal image-text representation learning
Chen, Y.-C., Li, L., Yu, L., El Kholy, A., Ahmed, F., Gan, Z., Cheng, Y., and Liu, J · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Cited alongside, same era.
Mule: Multimodal universal language embedding
Kim, D., Saito, K., Saenko, K., Sclaroff, S., and Plummer, B · 2020
Cited alongside, same era.
Jia, C., Yang, Y., Xia, Y., Chen, Y.-T., Parekh, Z., Pham, H., Le, Q., Sung, Y.-H., Li, Z., and Duerig, T · 2021
Later among the works it cites.
Vilt: Vision-and-language transformer without convolution or region supervision
Kim, W., Son, B., and Kim, I · 2021
Later among the works it cites.
Clip4clip: An empirical study of clip for end to end video clip retrieval
Luo, H., Ji, L., Zhong, M., Chen, Y., Lei, W., Duan, N., and Li, T · 2021
Later among the works it cites.
M3p: Learning universal representations via multitask multilingual multimodal pre-training
Ni, M., Huang, H., Su, L., Cui, E., Bharti, T., Wang, L., Zhang, D., and Duan, N · 2021
Later among the works it cites.
xgqa: Cross-lingual visual question answering
Pfeiffer, J., Geigle, G., Kamath, A., Steitz, J.-M. O., Roth, S., Vulić, I., and Gurevych, I · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Later among the works it cites.
Product-oriented machine translation with cross-modal cross-lingual pre-training
Song, Y., Chen, S., Jin, Q., Luo, W., Xie, J., and Huang, F · 2021
Later among the works it cites.
Florence: A new foundation model for computer vision
Yuan, L., Chen, D., Chen, Y.-L., Codella, N., Dai, X., Gao, J., Hu, H., Huang, X., Li, B., Li, C., et al · 2021
Later among the works it cites.
Uc2: Universal cross-lingual cross-modal vision-and-language pre-training
Zhou, M., Zhou, L., Wang, S., Cheng, Y., Li, L., Yu, Z., and Liu, J · 2021
Later among the works it cites.