Fetching the paper…
Reading the bibliography…
Recent advances in vision-language pre-training (VLP) have demonstrated impressive performance in a range of vision-language (VL) tasks.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 1907
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 1910
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Pixel-bert: Aligning image pixels with text by deep multi-modal transformers
Huang, Z., Zeng, Z., Liu, B., Fu, D., and Fu, J · 2004
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, W. B. and Brockett, C · 2005
Earlier work this paper cites.
The PASCAL recognising textual entailment challenge
Dagan, I., Glickman, O., and Magnini, B · 2006
Earlier work this paper cites.
Proceedings of the Fourth International Workshop on Semantic Evaluations (SemEval-2007) , Prague, Czech Republic, 2007. Association for Computational Linguistics
Agirre, E., Màrquez, L., and Wicentowski, R. (eds.) · 2007
Earlier work this paper cites.
Mind as machine: A history of cognitive science
Boden, M. A · 2008
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Ordonez, V., Kulkarni, G., and Berg, T. L · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C · 2013
Earlier work this paper cites.
Microsoft COCO: common objects in context
Lin, T., Maire, M., Belongie, S. J., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Earlier work this paper cites.
VQA: visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D · 2015
Earlier work this paper cites.
Microsoft COCO captions: Data collection and evaluation server
Chen, X., Fang, H., Lin, T., Vedantam, R., Gupta, S., Dollár, P., and Zitnick, C. L · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J., and Lazebnik, S · 2015
Earlier work this paper cites.
Faster R-CNN: towards real-time object detection with region proposal networks
Ren, S., He, K., Girshick, R. B., and Sun, J · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Vedantam, R., Zitnick, C. L., and Parikh, D · 2015
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Yu, L., Poirson, P., Yang, S., Berg, A. C., and Berg, T. L · 2016
Earlier work this paper cites.
Visual7w: Grounded question answering in images
Zhu, Y., Groth, O., Bernstein, M. S., and Fei-Fei, L · 2016
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
Johnson, J., Hariharan, B., van der Maaten, L., Fei-Fei, L., Zitnick, C. L., and Girshick, R. B · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L., Shamma, D. A., Bernstein, M. S., and Fei-Fei, L · 2017
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., and Zhang, L · 2018
Earlier work this paper cites.
Deep contextualized word representations
Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Unified language model pre-training for natural language understanding and generation
Dong, L., Yang, N., Wang, W., Wei, F., Liu, X., Wang, Y., Gao, J., Zhou, M., and Hon, H · 2019
Cited alongside, same era.
GQA: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A. and Manning, C. D · 2019
Cited alongside, same era.
Visualbert: A simple and performant baseline for vision and language
Li, L. H., Yatskar, M., Yin, D., Hsieh, C., and Chang, K · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J., Batra, D., Parikh, D., and Lee, S · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Cited alongside, same era.
BERT loses patience: Fast and robust inference with early exit
Zhou, W., Xu, C., Ge, T., McAuley, J. J., Xu, K., and Wei, F · 2020
Later among the works it cites.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Changpinyo, S., Sharma, P., Ding, N., and Soricut, R · 2021
Later among the works it cites.
Unifying vision-and-language tasks via text generation
Cho, J., Lei, J., Tan, H., and Bansal, M · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Later among the works it cites.
An empirical study of training end-to-end vision-and-language transformers
Dou, Z., Xu, Y., Gan, Z., Wang, J., Wang, S., Wang, L., Zhu, C., Zhang, P., Yuan, L., Peng, N., Liu, Z., and Zeng, M · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Cited alongside, same era.
A corpus for reasoning about natural language grounded in photographs
Suhr, A., Zhou, S., Zhang, A., Zhang, I., Bai, H., and Artzi, Y · 2019
Cited alongside, same era.
LXMERT: Learning cross-modality encoder representations from transformers
Tan, H. and Bansal, M · 2019
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Cited alongside, same era.
Neural network acceptability judgments
Warstadt, A., Singh, A., and Bowman, S. R · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Seeing out of the box: End-to-end pre-training for vision-language representation learning
Huang, Z., Zeng, Z., Huang, Y., Liu, B., Fu, D., and Fu, J · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Jia, C., Yang, Y., Xia, Y., Chen, Y., Parekh, Z., Pham, H., Le, Q. V., Sung, Y., Li, Z., and Duerig, T · 2021
Later among the works it cites.
Mdetr-modulated detection for end-to-end multi-modal understanding
Kamath, A., Singh, M., LeCun, Y., Synnaeve, G., Misra, I., and Carion, N · 2021
Later among the works it cites.
Vilt: Vision-and-language transformer without convolution or region supervision
Kim, W., Son, B., and Kim, I · 2021
Later among the works it cites.
WILDS: A benchmark of in-the-wild distribution shifts
Koh, P. W., Sagawa, S., Marklund, H., Xie, S. M., Zhang, M., Balsubramani, A., Hu, W., Yasunaga, M., Phillips, R. L., Gao, I., Lee, T., David, E., Stavness, I., Guo, W., Earnshaw, B., Haque, I., Beery, S. M., Leskovec, J., Kundaje, A., Pierson, E., Levine, S., Finn, C., and Liang, P · 2021
Later among the works it cites.
Value: A multi-task benchmark for video-and-language understanding evaluation
Li, L., Lei, J., Gan, Z., Yu, L., Chen, Y.-C., Pillai, R., Cheng, Y., Zhou, L., Wang, X. E., Wang, W. Y., et al · 2021
Later among the works it cites.
UNIMO: Towards unified-modal understanding and generation via cross-modal contrastive learning
Li, W., Gao, C., Niu, G., Xiao, X., Liu, H., Liu, J., Wu, H., and Wang, H · 2021
Later among the works it cites.
Visually grounded reasoning across languages and cultures
Liu, F., Bugliarello, E., Ponti, E. M., Reddy, S., Collier, N., and Elliott, D · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Later among the works it cites.
GEM: A general evaluation benchmark for multimodal tasks
Su, L., Duan, N., Cui, E., Ji, L., Wu, C., Luo, H., Liu, Y., Zhong, M., Bharti, T., and Sacheti, A · 2021
Later among the works it cites.
Long range arena : A benchmark for efficient transformers
Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D · 2021
Later among the works it cites.
Beyond preserved accuracy: Evaluating loyalty and robustness of BERT compression
Xu, C., Zhou, W., Ge, T., Xu, K., McAuley, J., and Wei, F · 2021
Later among the works it cites.
Beyond preserved accuracy: Evaluating loyalty and robustness of BERT compression
Xu, C., Zhou, W., Ge, T., Xu, K., McAuley, J., and Wei, F · 2021
Later among the works it cites.
Blow the dog whistle: A Chinese dataset for cant understanding with common sense and world knowledge
Xu, C., Zhou, W., Ge, T., Xu, K., McAuley, J., and Wei, F · 2021
Later among the works it cites.
E2E-VLP: End-to-end vision-language pre-training enhanced by visual learning
Xu, H., Yan, M., Li, C., Bi, B., Huang, S., Xiao, W., and Huang, F · 2021
Later among the works it cites.
Crossing the format boundary of text and boxes: Towards unified vision-language modeling
Yang, Z., Gan, Z., Wang, J., Hu, X., Ahmed, F., Liu, Z., Lu, Y., and Wang, L · 2021
Later among the works it cites.
Multi-grained vision language pre-training: Aligning texts with visual concepts
Zeng, Y., Zhang, X., and Li, H · 2021
Later among the works it cites.
Vinvl: Revisiting visual representations in vision-language models
Zhang, P., Li, X., Hu, X., Yang, J., Zhang, L., Wang, L., Choi, Y., and Gao, J · 2021
Later among the works it cites.
Improving sequence-to-sequence pre-training via sequence span rewriting
Zhou, W., Ge, T., Xu, C., Xu, K., and Wei, F · 2021
Later among the works it cites.
Pre-training text-to-text transformers for concept-centric common sense
Zhou, W., Lee, D., Selvam, R. K., Lee, S., and Ren, X · 2021
Later among the works it cites.
Contextual representation learning beyond masked language modeling
Fu, Z., Zhou, W., Xu, J., Zhou, H., and Li, L · 2022
Closest in time.
Wang, P., Yang, A., Men, R., Lin, J., Bai, S., Li, Z., Ma, J., Zhou, C., Zhou, J., and Yang, H · 2022
Closest in time.
A corpus of natural language for visual reasoning
Suhr, A., Lewis, M., Yeh, J., and Artzi, Y · 2034
Closest in time.