Fetching the paper…
Reading the bibliography…
The cost of vision-and-language pre-training has become increasingly prohibitive due to end-to-end training of large-scale models.
Im2text: Describing images using 1 million captioned photographs
Ordonez, V., Kulkarni, G., and Berg, T. L · 2011
Earlier work this paper cites.
Microsoft COCO: common objects in context
Lin, T., Maire, M., Belongie, S. J., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J., and Lazebnik, S · 2015
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L., Shamma, D. A., Bernstein, M. S., and Fei-Fei, L · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R · 2018
Earlier work this paper cites.
nocaps: novel object captioning at scale
Agrawal, H., Anderson, P., Desai, K., Wang, Y., Chen, X., Jain, R., Johnson, M., Batra, D., Parikh, D., and Lee, S · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Unified language model pre-training for natural language understanding and generation
Dong, L., Yang, N., Wang, W., Wei, F., Liu, X., Wang, Y., Gao, J., Zhou, M., and Hon, H · 2019
Earlier work this paper cites.
GQA: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A. and Manning, C. D · 2019
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Marino, K., Rastegari, M., Farhadi, A., and Mottaghi, R · 2019
Earlier work this paper cites.
LXMERT: learning cross-modality encoder representations from transformers
Tan, H. and Bansal, M · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
UNITER: universal image-text representation learning
Chen, Y., Li, L., Yu, L., Kholy, A. E., Ahmed, F., Gan, Z., Cheng, Y., and Liu, J · 2020
Cited alongside, same era.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Li, X., Yin, X., Li, C., Zhang, P., Hu, X., Zhang, L., Wang, L., Hu, H., Dong, L., Wei, F., Choi, Y., and Gao, J · 2020
Cited alongside, same era.
Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Changpinyo, S., Sharma, P., Ding, N., and Soricut, R · 2021
Cited alongside, same era.
Unifying vision-and-language tasks via text generation
Cho, J., Lei, J., Tan, H., and Bansal, M · 2021
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., Webson, A., Gu, S. S., Dai, Z., Suzgun, M., Chen, X., Chowdhery, A., Narang, S., Mishra, G., Yu, A., Zhao, V. Y., Huang, Y., Dai, A. M., Yu, H., Petrov, S., Chi, E. H., Dean, J., Devlin, J., Roberts, A., Zhou, D., Le, Q. V., and Wei, J · 2022
Later among the works it cites.
Enabling multimodal generation on CLIP via vision-language knowledge distillation
Dai, W., Hou, L., Shang, L., Jiang, X., Liu, Q., and Fung, P · 2022
Later among the works it cites.
Eva: Exploring the limits of masked visual representation learning at scale
Fang, Y., Wang, W., Xie, B., Sun, Q., Wu, L., Wang, X., Huang, T., Wang, X., and Cao, Y · 2022
Later among the works it cites.
From images to textual prompts: Zero-shot VQA with frozen large language models
Guo, J., Li, J., Li, D., Tiong, A. M. H., Li, B., Tao, D., and Hoi, S. C. H · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
Jia, C., Yang, Y., Xia, Y., Chen, Y.-T., Parekh, Z., Pham, H., Le, Q. V., Sung, Y., Li, Z., and Duerig, T · 2021
Cited alongside, same era.
Align before fuse: Vision and language representation learning with momentum distillation
Li, J., Selvaraju, R. R., Gotmare, A. D., Joty, S., Xiong, C., and Hoi, S · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Schuhmann, C., Vencu, R., Beaumont, R., Kaczmarczyk, R., Mullis, C., Katta, A., Coombes, T., Jitsev, J., and Komatsuzaki, A · 2021
Cited alongside, same era.
Multimodal few-shot learning with frozen language models
Tsimpoukelli, M., Menick, J., Cabi, S., Eslami, S. M. A., Vinyals, O., and Hill, F · 2021
Cited alongside, same era.
Florence: A new foundation model for computer vision
Yuan, L., Chen, D., Chen, Y., Codella, N., Dai, X., Gao, J., Hu, H., Huang, X., Li, B., Li, C., Liu, C., Liu, M., Liu, Z., Lu, Y., Shi, Y., Wang, L., Wang, J., Xiao, B., Xiao, Z., Yang, J., Zeng, M., Zhou, L., and Zhang, P · 2021
Cited alongside, same era.
Vinvl: Making visual representations matter in vision-language models
Zhang, P., Li, X., Hu, X., Yang, J., Zhang, L., Wang, L., Choi, Y., and Gao, J · 2021
Cited alongside, same era.
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., Driessche, G. v. d., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J. W., Vinyals, O., and Sifre, L · 2022
Later among the works it cites.
A good prompt is worth millions of parameters: Low-resource prompt-based learning for vision-language models
Jin, W., Cheng, Y., Shen, Y., Chen, W., and Ren, X · 2022
Later among the works it cites.
BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J., Li, D., Xiong, C., and Hoi, S. C. H · 2022
Later among the works it cites.
Plug-and-play VQA: zero-shot VQA by conjoining large pretrained models with zero training
Tiong, A. M. H., Li, J., Li, B., Savarese, S., and Hoi, S. C. H · 2022
Later among the works it cites.
FILIP: fine-grained interactive language-image pre-training
Yao, L., Huang, R., Hou, L., Lu, G., Niu, M., Xu, H., Liang, X., Li, Z., Jiang, X., and Xu, C · 2022
Later among the works it cites.
Coca: Contrastive captioners are image-text foundation models
Yu, J., Wang, Z., Vasudevan, V., Yeung, L., Seyedhosseini, M., and Wu, Y · 2022
Later among the works it cites.
Lit: Zero-shot transfer with locked-image text tuning
Zhai, X., Wang, X., Mustafa, B., Steiner, A., Keysers, D., Kolesnikov, A., and Beyer, L · 2022
Later among the works it cites.
OPT: open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M. T., Li, X., Lin, X. V., Mihaylov, T., Ott, M., Shleifer, S., Shuster, K., Simig, D., Koura, P. S., Sridhar, A., Wang, T., and Zettlemoyer, L · 2022
Later among the works it cites.
MAPL: parameter-efficient adaptation of unimodal pre-trained models for vision-language few-shot prompting
Mañas, O., Rodríguez, P., Ahmadi, S., Nematzadeh, A., Goyal, Y., and Agrawal, A · 2023
Closest in time.