Fetching the paper…
Reading the bibliography…
This paper surveys vision-language pre-training (VLP) methods for multimodal intelligence that have been developed in the last few years.
Visual entailment: A novel task for fine-grained image understanding
Xie, N., Lai, F., Doran, D., and Kadav, A. (2019) · 1901
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019d) · 1907
Earlier work this paper cites.
VisualBERT: A simple and performant baseline for vision and language
Li, L. H., Yatskar, M., Yin, D., Hsieh, C.-J., and Chang, K.-W. (2019e) · 1908
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B. (2019) · 1909
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Univl: A unified video and language pre-training model for multimodal understanding and generation
Luo, H., Ji, L., Shi, B., Huang, H., Duan, N., Li, T., Li, J., Bharti, T., and Zhou, M. (2020) · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. (2002) · 2002
Earlier work this paper cites.
Pixel-BERT: Aligning image pixels with text by deep multi-modal transformers
Huang, Z., Zeng, Z., Liu, B., Fu, D., and Fu, J. (2020) · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y. (2004) · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Banerjee, S. and Lavie, A. (2005) · 2005
Earlier work this paper cites.
Auto-captions on gif: A large-scale video-sentence dataset for vision-language pre-training
Pan, Y., Li, Y., Luo, J., Xu, J., Yao, T., and Mei, T. (2020a) · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009) · 2009
Earlier work this paper cites.
Describing objects by their attributes
Farhadi, A., Endres, I., Hoiem, D., and Forsyth, D. (2009) · 2009
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
Farhadi, A., Hejrati, M., Sadeghi, M. A., Young, P., Rashtchian, C., Hockenmaier, J., and Forsyth, D. (2010) · 2010
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M. and Hyvärinen, A. (2010) · 2010
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
Chen, D. and Dolan, W. (2011) · 2011
Earlier work this paper cites.
HMDB: a large video database for human motion recognition
Kuehne, H., Jhuang, H., Garrote, E., Poggio, T., and Serre, T. (2011) · 2011
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Ordonez, V., Kulkarni, G., and Berg, T. (2011) · 2011
Earlier work this paper cites.
Evaluating knowledge transfer and zero-shot learning in a large-scale setting
Rohrbach, M., Stark, M., and Schiele, B. (2011) · 2011
Earlier work this paper cites.
The Caltech-UCSD birds-200-2011 dataset
Wah, C., Branson, S., Welinder, P., Perona, P., and Belongie, S. (2011) · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012) · 2012
Earlier work this paper cites.
A closer look at the robustness of vision-and-language pre-trained models
Li, L., Gan, Z., and Liu, J. (2020c) · 2012
Earlier work this paper cites.
Seeing past words: Testing the cross-modal capabilities of pretrained v&l models on counting tasks
Parcalabescu, L., Gatt, A., Frank, A., and Calixto, I. (2020) · 2012
Earlier work this paper cites.
SUN attribute database: Discovering, annotating, and recognizing scene attributes
Patterson, G. and Hays, J. (2012) · 2012
Earlier work this paper cites.
Minivlm: A smaller and faster vision-language model
Wang, J., Hu, X., Zhang, P., Li, X., Wang, L., Zhang, L., Gao, J., and Liu, Z. (2020a) · 2012
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Frome, A., Corrado, G. S., Shlens, J., Bengio, S., Dean, J., Ranzato, M., and Mikolov, T. (2013) · 2013
Earlier work this paper cites.
Babytalk: Understanding and generating simple image descriptions
Kulkarni, G., Premraj, V., Ordonez, V., Dhar, S., Li, S., Choi, Y., Berg, A. C., and Berg, T. L. (2013) · 2013
Earlier work this paper cites.
Attribute-based classification for zero-shot visual object categorization
Lampert, C. H., Nickisch, H., and Harmeling, S. (2013) · 2013
Earlier work this paper cites.
Grounding action descriptions in videos
Regneri, M., Rohrbach, M., Wetzel, D., Thater, S., Schiele, B., and Pinkal, M. (2013) · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Cho, K., van Merrienboer, B., Gülçehre, Ç., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014) · 2014
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
Kiros, R., Salakhutdinov, R., and Zemel, R. S. (2014) · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L. (2014) · 2014
Earlier work this paper cites.
Deep captioning with multimodal recurrent neural networks (m-rnn)
Mao, J., Xu, W., Yang, Y., Wang, J., Huang, Z., and Yuille, A. (2014) · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C. D. (2014) · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A. (2014) · 2014
Earlier work this paper cites.
Grounded compositional semantics for finding and describing images with sentences
Socher, R., Karpathy, A., Le, Q. V., Manning, C. D., and Ng, A. Y. (2014) · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V. (2014) · 2014
Earlier work this paper cites.
Edge boxes: Locating object proposals from edges
Zitnick, C. L. and Dollár, P. (2014) · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D. (2015) · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y. (2015) · 2015
Earlier work this paper cites.
Microsoft COCO captions: Data collection and evaluation server
Chen, X., Fang, H., Lin, T., Vedantam, R., Gupta, S., Dollár, P., and Zitnick, C. L. (2015) · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Donahue, J., Anne Hendricks, L., Guadarrama, S., Rohrbach, M., Venugopalan, S., Saenko, K., and Darrell, T. (2015) · 2015
Earlier work this paper cites.
From captions to visual concepts and back
Fang, H., Gupta, S., Iandola, F., Srivastava, R. K., Deng, L., Dollár, P., Gao, J., He, X., Mitchell, M., Platt, J. C., et al. (2015) · 2015
Earlier work this paper cites.
Are you talking to a machine? dataset and methods for multilingual image question
Gao, H., Mao, J., Zhou, J., Huang, Z., Wang, L., and Xu, W. (2015) · 2015
Earlier work this paper cites.
Fast r-cnn
Girshick, R. (2015) · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A. and Fei-Fei, L. (2015) · 2015
Earlier work this paper cites.
Associating neural word embeddings with deep image representations using fisher vectors
Klein, B., Lev, G., Sadeh, G., and Wolf, L. (2015) · 2015
Earlier work this paper cites.
Ask your neurons: A neural-based approach to answering questions about images
Malinowski, M., Rohrbach, M., and Fritz, M. (2015) · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J., and Lazebnik, S. (2015) · 2015
Earlier work this paper cites.
A Dataset for Movie Description
Rohrbach, A., Rohrbach, M., Tandon, N., and Schiele, B. (2015) · 2015
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015) · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M. (2015) · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Vedantam, R., Lawrence Zitnick, C., and Parikh, D. (2015) · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Vinyals, O., Toshev, A., Bengio, S., and Erhan, D. (2015) · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., and Bengio, Y. (2015) · 2015
Earlier work this paper cites.
Visual madlibs: Fill in the blank image generation and question answering
Yu, L., Park, E., Berg, A. C., and Berg, T. L. (2015) · 2015
Earlier work this paper cites.
Youtube-8m: A large-scale video classification benchmark
Abu-El-Haija, S., Kothari, N., Lee, J., Natsev, P., Toderici, G., Varadarajan, B., and Vijayanarasimhan, S. (2016) · 2016
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
Anderson, P., Fernando, B., Johnson, M., and Gould, S. (2016) · 2016
Earlier work this paper cites.
Semi-supervised vocabulary-informed learning
Fu, Y. and Sigal, L. (2016) · 2016
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Fukui, A., Park, D. H., Yang, D., Rohrbach, A., Darrell, T., and Rohrbach, M. (2016) · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Earlier work this paper cites.
Visual storytelling
Huang, T.-H., Ferraro, F., Mostafazadeh, N., Misra, I., Agrawal, A., Devlin, J., Girshick, R., He, X., Kohli, P., Batra, D., et al. (2016) · 2016
Earlier work this paper cites.
Revisiting visual question answering baselines
Jabri, A., Joulin, A., and Maaten, L. v. d. (2016) · 2016
Earlier work this paper cites.
Exploring the limits of language modeling
Jozefowicz, R., Vinyals, O., Schuster, M., Shazeer, N., and Wu, Y. (2016) · 2016
Earlier work this paper cites.
Multimodal residual learning for visual qa
Kim, J.-H., Lee, S.-W., Kwak, D., Heo, M.-O., Kim, J., Ha, J.-W., and Zhang, B.-T. (2016) · 2016
Earlier work this paper cites.
Hierarchical question-image co-attention for visual question answering
Lu, J., Yang, J., Batra, D., and Parikh, D. (2016) · 2016
Earlier work this paper cites.
Learning to answer questions from image using convolutional neural network
Ma, L., Lu, Z., and Li, H. (2016) · 2016
Earlier work this paper cites.
Generating images from captions with attention
Mansimov, E., Parisotto, E., Ba, J. L., and Salakhutdinov, R. (2016) · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Mao, J., Huang, J., Toshev, A., Camburu, O., Yuille, A. L., and Murphy, K. (2016) · 2016
Earlier work this paper cites.
Modeling context between objects for referring expression understanding
Nagaraja, V. K., Morariu, V. I., and Davis, L. S. (2016) · 2016
Earlier work this paper cites.
A benchmark dataset and evaluation methodology for video object segmentation
Perazzi, F., Pont-Tuset, J., McWilliams, B., Van Gool, L., Gross, M., and Sorkine-Hornung, A. (2016) · 2016
Earlier work this paper cites.
Generative adversarial text to image synthesis
Reed, S., Akata, Z., Yan, X., Logeswaran, L., Schiele, B., and Lee, H. (2016) · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A. (2016) · 2016
Earlier work this paper cites.
Where to look: Focus regions for visual question answering
Shih, K. J., Singh, S., and Hoiem, D. (2016) · 2016
Earlier work this paper cites.
Learning Language-Visual Embedding for Movie Understanding with Natural-Language
Torabi, A., Tandon, N., and Sigal, L. (2016) · 2016
Earlier work this paper cites.
Learning deep structure-preserving image-text embeddings
Wang, L., Li, Y., and Lazebnik, S. (2016) · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., et al. (2016) · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Xu, J., Mei, T., Yao, T., and Rui, Y. (2016) · 2016
Earlier work this paper cites.
Stacked attention networks for image question answering
Yang, Z., He, X., Gao, J., Deng, L., and Smola, A. (2016) · 2016
Earlier work this paper cites.
Image captioning with semantic attention
You, Q., Jin, H., Wang, Z., Fang, C., and Luo, J. (2016) · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Yu, L., Poirson, P., Yang, S., Berg, A. C., and Berg, T. L. (2016) · 2016
Earlier work this paper cites.
Visual7w: Grounded question answering in images
Zhu, Y., Groth, O., Bernstein, M., and Fei-Fei, L. (2016) · 2016
Earlier work this paper cites.
Mutan: Multimodal tucker fusion for visual question answering
Ben-Younes, H., Cadene, R., Cord, M., and Thome, N. (2017) · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J. and Zisserman, A. (2017) · 2017
Earlier work this paper cites.
Visual dialog
Das, A., Kottur, S., Gupta, K., Singh, A., Yadav, D., Moura, J. M., Parikh, D., and Batra, D. (2017) · 2017
Earlier work this paper cites.
Vse++: Improving visual-semantic embeddings with hard negatives
Faghri, F., Fleet, D. J., Kiros, J. R., and Fidler, S. (2017) · 2017
Earlier work this paper cites.
Localizing Moments in Video with Natural Language
Hendricks, L. A., Wang, O., Shechtman, E., Sivic, J., Darrell, T., and Russell, B. (2017) · 2017
Earlier work this paper cites.
Learning to reason: End-to-end module networks for visual question answering
Hu, R., Andreas, J., Rohrbach, M., Darrell, T., and Saenko, K. (2017) · 2017
Earlier work this paper cites.
Instance-aware image and sentence matching with selective multimodal lstm
Huang, Y., Wang, W., and Wang, L. (2017) · 2017
Earlier work this paper cites.
TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering
Jang, Y., Song, Y., Yu, Y., Kim, Y., and Kim, G. (2017) · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., et al. (2017) · 2017
Earlier work this paper cites.
Hadamard product for low-rank bilinear pooling
Kim, J.-H., On, K.-W., Lim, W., Kim, J., Ha, J.-W., and Zhang, B.-T. (2017) · 2017
Earlier work this paper cites.
A hierarchical approach for generating descriptive image paragraphs
Krause, J., Johnson, J., Krishna, R., and Fei-Fei, L. (2017) · 2017
Earlier work this paper cites.
Knowing when to look: Adaptive attention via a visual sentinel for image captioning
Lu, J., Xiong, C., Parikh, D., and Socher, R. (2017) · 2017
Earlier work this paper cites.
Dual attention networks for multimodal reasoning and matching
Nam, H., Ha, J.-W., and Kim, J. (2017) · 2017
Earlier work this paper cites.
Hierarchical multimodal lstm for dense visual-semantic embedding
Niu, Z., Zhou, M., Wang, L., Gao, X., and Hua, G. (2017) · 2017
Earlier work this paper cites.
A simple neural network module for relational reasoning
Santoro, A., Raposo, D., Barrett, D. G., Malinowski, M., Pascanu, R., Battaglia, P., and Lillicrap, T. (2017) · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J. (2017) · 2017
Earlier work this paper cites.
Tips and tricks for visual question answering: Learnings from the 2017 challenge
Teney, D., Anderson, P., He, X., and van den Hengel, A. (2018) · 2017
Earlier work this paper cites.
Graph-structured representations for visual question answering
Teney, D., Liu, L., and van Den Hengel, A. (2017) · 2017
Earlier work this paper cites.
Neural discrete representation learning
van den Oord, A., Vinyals, O., and Kavukcuoglu, K. (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Video Question Answering via Gradually Refined Attention over Appearance and Motion
Xu, D., Zhao, Z., Xiao, J., Wu, F., Zhang, H., He, X., and Zhuang, Y. (2017) · 2017
Earlier work this paper cites.
Boosting image captioning with attributes
Yao, T., Pan, Y., Li, Y., Qiu, Z., and Mei, T. (2017) · 2017
Earlier work this paper cites.
Multi-modal factorized bilinear pooling with co-attention learning for visual question answering
Yu, Z., Yu, J., Fan, J., and Tao, D. (2017) · 2017
Earlier work this paper cites.
Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks
Zhang, H., Xu, T., Li, H., Zhang, S., Wang, X., Huang, X., and Metaxas, D. N. (2017) · 2017
Earlier work this paper cites.
Don’t just assume; look and answer: Overcoming priors for visual question answering
Agrawal, A., Batra, D., Parikh, D., and Kembhavi, A. (2018) · 2018
Earlier work this paper cites.
Convolutional image captioning
Aneja, J., Deshpande, A., and Schwing, A. G. (2018) · 2018
Earlier work this paper cites.
Real-time referring expression comprehension by single-stage grounding network
Chen, X., Ma, L., Chen, J., Jie, Z., Liu, W., and Luo, J. (2018) · 2018
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
Damen, D., Doughty, H., Farinella, G. M., Fidler, S., Furnari, A., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., et al. (2018) · 2018
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
Gurari, D., Li, Q., Stangl, A. J., Guo, A., Lin, C., Grauman, K., Luo, J., and Bigham, J. P. (2018) · 2018
Earlier work this paper cites.
Explainable neural computation via stack neural module networks
Hu, R., Andreas, J., Darrell, T., and Saenko, K. (2018) · 2018
Earlier work this paper cites.
Compositional attention networks for machine reasoning
Hudson, D. A. and Manning, C. D. (2018) · 2018
Earlier work this paper cites.
Bilinear attention networks
Kim, J.-H., Jun, J., and Zhang, B.-T. (2018) · 2018
Earlier work this paper cites.
Stacked cross attention for image-text matching
Lee, K.-H., Chen, X., Hua, G., Hu, H., and He, X. (2018) · 2018
Earlier work this paper cites.
Tvqa: Localized, compositional video question answering
Lei, J., Yu, L., Bansal, M., and Berg, T. L. (2018) · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F. (2018) · 2018
Earlier work this paper cites.
Responsible bots: 10 guidelines for developers of conversational ai
Microsoft (2018) · 2018
Earlier work this paper cites.
Improved fusion of visual and language representations by dense symmetric co-attention for visual question answering
Nguyen, D.-K. and Okatani, T. (2018) · 2018
Cited alongside, same era.
Learning conditioned graph structures for interpretable visual question answering
Norcliffe-Brown, W., Vafeias, S., and Parisot, S. (2018) · 2018
Cited alongside, same era.
Film: Visual reasoning with a general conditioning layer
Perez, E., Strub, F., De Vries, H., Dumoulin, V., and Courville, A. (2018) · 2018
Cited alongside, same era.
Conditional image-text embedding networks
Plummer, B. A., Kordas, P., Kiapour, M. H., Zheng, S., Piramuthu, R., and Lazebnik, S. (2018) · 2018
Cited alongside, same era.
Yolov3: An incremental improvement
Redmon, J. and Farhadi, A. (2018) · 2018
Cited alongside, same era.
Krisp: Integrating implicit and symbolic knowledge for open-domain knowledge-based vqa
Marino, K., Chen, X., Parikh, D., Gupta, A., and Rohrbach, M. (2021) · 2021
Later among the works it cites.
Deep learning–based text classification: a comprehensive review
Minaee, S., Kalchbrenner, N., Cambria, E., Nikzad, N., Chenaghlu, M., and Gao, J. (2021) · 2021
Later among the works it cites.
Slip: Self-supervision meets language-image pre-training
Mu, N., Kirillov, A., Wagner, D., and Xie, S. (2021) · 2021
Later among the works it cites.
M3p: Learning universal representations via multitask multilingual multimodal pre-training
Ni, M., Huang, H., Su, L., Cui, E., Bharti, T., Wang, L., Gao, J., Zhang, D., and Duan, N. (2021) · 2021
Later among the works it cites.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., and Chen, M. (2021) · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R. (2018) · 2018
Cited alongside, same era.
Learning two-branch neural networks for image-text matching tasks
Wang, L., Li, Y., Huang, J., and Lazebnik, S. (2018) · 2018
Cited alongside, same era.
Attngan: Fine-grained text to image generation with attentional generative adversarial networks
Xu, T., Zhang, P., Huang, Q., Zhang, H., Gan, Z., Huang, X., and He, X. (2018) · 2018
Cited alongside, same era.
Exploring visual relationship for image captioning
Yao, T., Pan, Y., Li, Y., and Mei, T. (2018) · 2018
Cited alongside, same era.
Neural-symbolic vqa: Disentangling reasoning from vision and language understanding
Yi, K., Wu, J., Gan, C., Torralba, A., Kohli, P., and Tenenbaum, J. (2018) · 2018
Cited alongside, same era.
Towards Automatic Learning of Procedures from Web Instructional Videos
Zhou, L., Xu, C., and Corso, J. J. (2018) · 2018
Cited alongside, same era.
nocaps: novel object captioning at scale
Agrawal, H., Desai, K., Wang, Y., Chen, X., Jain, R., Johnson, M., Batra, D., Parikh, D., Lee, S., and Anderson, P. (2019) · 2019
Cited alongside, same era.
Later among the works it cites.
Mlp architectures for vision-and-language modeling: An empirical study
Nie, Y., Li, L., Gan, Z., Wang, S., Zhu, C., Zeng, M., Liu, Z., Bansal, M., and Wang, L. (2021) · 2021
Later among the works it cites.
Valse: A task-independent benchmark for vision and language models centered on linguistic phenomena
Parcalabescu, L., Cafagna, M., Muradjan, L., Frank, A., Calixto, I., and Gatt, A. (2021) · 2021
Later among the works it cites.
Combined scaling for zero-shot transfer learning
Pham, H., Dai, Z., Ghiasi, G., Liu, H., Yu, A. W., Luong, M.-T., Tan, M., and Le, Q. V. (2021) · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. (2021) · 2021
Later among the works it cites.
Zero-Shot Text-to-Image Generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I. (2021) · 2021
Later among the works it cites.
Vision transformers for dense prediction
Ranftl, R., Bochkovskiy, A., and Koltun, V. (2021) · 2021
Later among the works it cites.
Denseclip: Language-guided dense prediction with context-aware prompting
Rao, Y., Zhao, W., Chen, G., Tang, Y., Zhu, Z., Huang, G., Zhou, J., and Lu, J. (2021) · 2021
Later among the works it cites.
Avlnet: Learning audio-visual language representations from instructional videos
Rouditchenko, A., Boggust, A., Harwath, D., Chen, B., Joshi, D., Thomas, S., Audhkhasi, K., Kuehne, H., Panda, R., Feris, R., et al. (2021) · 2021
Later among the works it cites.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Schuhmann, C., Vencu, R., Beaumont, R., Kaczmarczyk, R., Mullis, C., Katta, A., Coombes, T., Jitsev, J., and Komatsuzaki, A. (2021) · 2021
Later among the works it cites.
Human-adversarial visual question answering
Sheng, S., Singh, A., Goswami, V., Magana, J., Thrush, T., Galuba, W., Parikh, D., and Kiela, D. (2021) · 2021
Later among the works it cites.
Reasoning over vision and language: Exploring the benefits of supplemental knowledge
Shevchenko, V., Teney, D., Dick, A., and Hengel, A. v. d. (2021) · 2021
Later among the works it cites.
Denoising diffusion implicit models
Song, J., Meng, C., and Ermon, S. (2021) · 2021
Later among the works it cites.
Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning
Srinivasan, K., Raman, K., Chen, J., Bendersky, M., and Najork, M. (2021) · 2021
Later among the works it cites.
Lightningdot: Pre-training visual-semantic embeddings for real-time image-text retrieval
Sun, S., Chen, Y.-C., Li, L., Wang, S., Fang, Y., and Liu, J. (2021) · 2021
Later among the works it cites.
VIMPAC: Video Pre-Training via Masked Token Prediction and Contrastive Learning
Tan, H., Lei, J., Wolf, T., and Bansal, M. (2021) · 2021
Later among the works it cites.
Vl-ltr: Learning class-wise visual-linguistic representation for long-tailed visual recognition
Tian, C., Wang, W., Zhu, X., Wang, X., Dai, J., and Qiao, Y. (2021) · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H. (2021) · 2021
Later among the works it cites.
Multimodal few-shot learning with frozen language models
Tsimpoukelli, M., Menick, J., Cabi, S., Eslami, S., Vinyals, O., and Hill, F. (2021) · 2021
Later among the works it cites.
On semantic similarity in video retrieval
Wray, M., Doughty, H., and Damen, D. (2021) · 2021
Later among the works it cites.
Star: A benchmark for situated reasoning in real-world videos
Wu, B., Yu, S., Chen, Z., Tenenbaum, J. B., and Gan, C. (2021) · 2021
Later among the works it cites.
Next-qa: Next phase of question-answering to explaining temporal actions
Xiao, J., Shang, X., Yao, A., and Chua, T.-S. (2021) · 2021
Later among the works it cites.
Probing inter-modality: Visual parsing with self-attention for vision-language pre-training
Xue, H., Huang, Y., Liu, B., Peng, H., Fu, J., Li, H., and Luo, J. (2021) · 2021
Later among the works it cites.
CPT: Colorful prompt tuning for pre-trained vision-language models
Yao, Y., Zhang, A., Zhang, Z., Liu, Z., Chua, T.-S., and Sun, M. (2021) · 2021
Later among the works it cites.
Ernie-vil: Knowledge enhanced vision-language representations through scene graphs
Yu, F., Tang, J., Yin, W., Sun, Y., Tian, H., Wu, H., and Wang, H. (2021) · 2021
Later among the works it cites.
Florence: A new foundation model for computer vision
Yuan, L., Chen, D., Chen, Y.-L., Codella, N., Dai, X., Gao, J., Hu, H., Huang, X., Li, B., Li, C., Liu, C., Liu, M., Liu, Z., Lu, Y., Shi, Y., Wang, L., Wang, J., Xiao, B., Xiao, Z., Yang, J., Zeng, M., Zhou, L., and Zhang, P. (2021) · 2021
Later among the works it cites.
Open-vocabulary object detection using captions
Zareian, A., Rosa, K. D., Hu, D. H., and Chang, S.-F. (2021) · 2021
Later among the works it cites.
Merlot: Multimodal neural script knowledge models
Zellers, R., Lu, X., Hessel, J., Yu, Y., Park, J. S., Cao, J., Farhadi, A., and Choi, Y. (2021) · 2021
Later among the works it cites.
Understanding and evaluating racial biases in image captioning
Zhao, D., Wang, A., and Russakovsky, O. (2021) · 2021
Later among the works it cites.
Uc2: Universal cross-lingual cross-modal vision-and-language pre-training
Zhou, M., Zhou, L., Wang, S., Cheng, Y., Li, L., Yu, Z., and Liu, J. (2021) · 2021
Later among the works it cites.
Kaleido-bert: Vision-language pre-training on fashion domain
Zhuge, M., Gao, D., Fan, D.-P., Jin, L., Chen, B., Zhou, H., Qiu, M., and Shao, L. (2021) · 2021
Later among the works it cites.
Cm3: A causal masked multimodal model of the internet
Aghajanyan, A., Huang, B., Ross, C., Karpukhin, V., Xu, H., Goyal, N., Okhonko, D., Joshi, M., Ghosh, G., Lewis, M., et al. (2022) · 2022
Closest in time.
Agrawal, A., Kajić, I., Bugliarello, E., Davoodi, E., Gergely, A., Blunsom, P., and Nematzadeh, A. (2022) · 2022
Closest in time.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al. (2022) · 2022
Closest in time.
Multimae: Multi-modal multi-task masked autoencoders
Bachmann, R., Mizrahi, D., Atanov, A., and Zamir, A. (2022) · 2022
Closest in time.
Latr: Layout-aware transformer for scene-text vqa
Biten, A. F., Litman, R., Xie, Y., Appalaraju, S., and Manmatha, R. (2022) · 2022
Closest in time.
Winogavil: Gamified association benchmark to challenge vision-and-language models
Bitton, Y., Guetta, N. B., Yosef, R., Elovici, Y., Bansal, M., Stanovsky, G., and Schwartz, R. (2022) · 2022
Closest in time.
Revisiting the” video” in video-language understanding
Buch, S., Eyzaguirre, C., Gaidon, A., Wu, J., Fei-Fei, L., and Niebles, J. C. (2022) · 2022
Closest in time.
X-detr: A versatile architecture for instance-wise vision-language tasks
Cai, Z., Kwon, G., Ravichandran, A., Bas, E., Tu, Z., Bhotika, R., and Soatto, S. (2022) · 2022
Closest in time.
Cross-lingual and multilingual clip
Carlsson, F., Eisen, P., Rekathati, F., and Sahlgren, M. (2022) · 2022
Closest in time.
Webqa: Multihop and multimodal qa
Chang, Y., Narang, M., Suzuki, H., Cao, G., Gao, J., and Bisk, Y. (2022) · 2022
Closest in time.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. (2022) · 2022
Closest in time.
One model, multiple modalities: A sparsely activated approach for text, sound, image, video and code
Dai, Y., Tang, D., Liu, L., Tan, M., Zhou, C., Wang, J., Feng, Z., Zhang, F., Hu, X., and Shi, S. (2022) · 2022
Closest in time.
Prefix language models are unified modal learners
Diao, S., Zhou, W., Zhang, X., and Wang, J. (2022) · 2022
Closest in time.
A survey of vision-language pre-trained models
Du, Y., Liu, Z., Li, J., and Zhao, W. X. (2022) · 2022
Closest in time.
Multi-modal alignment using representation codebook
Duan, J., Chen, L., Tran, S., Yang, J., Xu, Y., Zeng, B., and Chilimbi, T. (2022) · 2022
Closest in time.
Promptdet: Expand your detector vocabulary with uncurated images
Feng, C., Zhong, Y., Jie, Z., Chu, X., Ren, H., Wei, X., Xie, W., and Ma, L. (2022) · 2022
Closest in time.
An empirical study of end-to-end video-language transformers with masked visual modeling
Fu, T.-J., Li, L., Gan, Z., Lin, K., Wang, W. Y., Wang, L., and Liu, Z. (2022) · 2022
Closest in time.
Make-a-scene: Scene-based text-to-image generation with human priors
Gafni, O., Polyak, A., Ashual, O., Sheynin, S., Parikh, D., and Taigman, Y. (2022) · 2022
Closest in time.
Playing lottery tickets with vision and language
Gan, Z., Chen, Y.-C., Li, L., Chen, T., Cheng, Y., Wang, S., and Liu, J. (2022) · 2022
Closest in time.
Bridging video-text retrieval with multiple choice questions
Ge, Y., Ge, Y., Liu, X., Li, D., Shan, Y., Qie, X., and Luo, P. (2022) · 2022
Closest in time.
Multimodal masked autoencoders learn transferable representations
Geng, X., Liu, H., Lee, L., Schuurams, D., Levine, S., and Abbeel, P. (2022) · 2022
Closest in time.
Open-vocabulary image segmentation
Ghiasi, G., Gu, X., Cui, Y., and Lin, T.-Y. (2022) · 2022
Closest in time.
Cyclip: Cyclic contrastive language-image pretraining
Goel, S., Bansal, H., Bhatia, S., Rossi, R. A., Vinay, V., and Grover, A. (2022) · 2022
Closest in time.
Fashionvlp: Vision language transformer for fashion retrieval with feedback
Goenka, S., Zheng, Z., Jaiswal, A., Chada, R., Wu, Y., Hedau, V., and Natarajan, P. (2022) · 2022
Closest in time.
When, why, and which pretrained gans are useful?
Grigoryev, T., Voynov, A., and Babenko, A. (2022) · 2022
Closest in time.
Kat: A knowledge augmented transformer for vision-and-language
Gui, L., Wang, B., Huang, Q., Hauptmann, A., Bisk, Y., and Gao, J. (2022) · 2022
Closest in time.
Temporal alignment networks for long-term video
Han, T., Xie, W., and Zisserman, A. (2022) · 2022
Closest in time.
Language models are general-purpose interfaces
Hao, Y., Song, H., Dong, L., Huang, S., Chi, Z., Wang, W., Ma, S., and Wei, F. (2022) · 2022
Closest in time.
Parameter-efficient fine-tuning for vision transformers
He, X., Li, C., Zhang, P., Yang, J., and Wang, X. E. (2022) · 2022
Closest in time.
Imagen video: High definition video generation with diffusion models
Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., et al. (2022) · 2022
Closest in time.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al. (2022) · 2022
Closest in time.
Scaling up vision-language pre-training for image captioning
Hu, X., Gan, Z., Wang, J., Yang, Z., Liu, Z., Lu, Y., and Wang, L. (2022) · 2022
Closest in time.
Du-vlg: Unifying vision-and-language generation via dual sequence-to-sequence pre-training
Huang, L., Niu, G., Liu, J., Xiao, X., and Wu, H. (2022) · 2022
Closest in time.
Carets: A consistency and robustness evaluative test suite for vqa
Jimenez, C. E., Russakovsky, O., and Narasimhan, K. (2022) · 2022
Closest in time.
A good prompt is worth millions of parameters? low-resource prompt-based learning for vision-language models
Jin, W., Cheng, Y., Shen, Y., Chen, W., and Ren, X. (2022) · 2022
Closest in time.
Prompting visual-language models for efficient video understanding
Ju, C., Han, T., Zheng, K., Zhang, Y., and Xie, W. (2022) · 2022
Closest in time.
Webly supervised concept expansion for general purpose vision models
Kamath, A., Clark, C., Gupta, T., Kolve, E., Hoiem, D., and Kembhavi, A. (2022) · 2022
Closest in time.
L-verse: Bidirectional generation between image and text
Kim, T., Song, G., Lee, S., Kim, S., Seo, Y., Lee, S., Kim, S. H., Lee, H., and Bae, K. (2022) · 2022
Closest in time.
Large-scale bilingual language-image contrastive learning
Ko, B. and Gu, G. (2022) · 2022
Closest in time.
Uvim: A unified modeling approach for vision with learned guiding codes
Kolesnikov, A., Pinto, A. S., Beyer, L., Zhai, X., Harmsen, J., and Houlsby, N. (2022) · 2022
Closest in time.
Fine-tuning can distort pretrained features and underperform out-of-distribution
Kumar, A., Raghunathan, A., Jones, R., Ma, T., and Liang, P. (2022) · 2022
Closest in time.
Findit: Generalized localization with natural language queries
Kuo, W., Bertsch, F., Li, W., Piergiovanni, A., Saffar, M., and Angelova, A. (2022) · 2022
Closest in time.
Masked vision and language modeling for multi-modal representation learning
Kwon, G., Cai, Z., Ravichandran, A., Bas, E., Bhotika, R., and Soatto, S. (2022) · 2022
Closest in time.
Video Swin Transformer
Liu, Z., Ning, J., Cao, Y., Wei, Y., Zhang, Z., Lin, S., and Hu, H. (2022) · 2022
Closest in time.
Image segmentation using text and image prompts
Lüddecke, T. and Ecker, A. (2022) · 2022
Closest in time.
Simple open-vocabulary object detection with vision transformers
Minderer, M., Gritsenko, A., Stone, A., Neumann, M., Weissenborn, D., Dosovitskiy, A., Mahendran, A., Arnab, A., Dehghani, M., Shen, Z., et al. (2022) · 2022
Closest in time.
Multimodal contrastive learning with limoe: the language-image mixture of experts
Mustafa, B., Riquelme, C., Puigcerver, J., Jenatton, R., and Houlsby, N. (2022) · 2022
Closest in time.
Expanding language-image pretrained models for general video recognition
Ni, B., Peng, H., Chen, M., Zhang, S., Meng, G., Fu, J., Xiang, S., and Ling, H. (2022) · 2022
Closest in time.
Exposing the limits of video-text models through contrast sets
Park, J. S., Shen, S., Farhadi, A., Darrell, T., Choi, Y., and Rohrbach, A. (2022) · 2022
Closest in time.
Data cards: Purposeful and transparent dataset documentation for responsible ai
Pushkarna, M., Zaldivar, A., and Kjartansson, O. (2022) · 2022
Closest in time.
Multimodal open-vocabulary video classification via pre-trained vision and language models
Qian, R., Li, Y., Xu, Z., Yang, M.-H., Belongie, S., and Cui, Y. (2022) · 2022
Closest in time.
Hierarchical text-conditional image generation with clip latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. (2022) · 2022
Closest in time.
Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., et al. (2022) · 2022
Closest in time.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. (2022) · 2022
Closest in time.
Survey: Transformer based video-language pre-training
Ruan, L. and Jin, Q. (2022) · 2022
Closest in time.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., et al. (2022) · 2022
Closest in time.
Prefix conditioning unifies language and label supervision
Saito, K., Sohn, K., Zhang, X., Li, C.-L., Lee, C.-Y., Saenko, K., and Pfister, T. (2022) · 2022
Closest in time.
Are vision-language transformers learning multimodal representations? a probing perspective
Salin, E., Farah, B., Ayache, S., and Favre, B. (2022) · 2022
Closest in time.
A-okvqa: A benchmark for visual question answering using world knowledge
Schwenk, D., Khandelwal, A., Clark, C., Marino, K., and Mottaghi, R. (2022) · 2022
Closest in time.
End-to-end generative pretraining for multimodal video captioning
Seo, P. H., Nagrani, A., Arnab, A., and Schmid, C. (2022) · 2022
Closest in time.
Ruclip–new models and experiments: a technical report
Shonenkov, A., Kuznetsov, A., Dimitrov, D., Shavrina, T., Chesakov, D., Maltseva, A., Fenogenova, A., Pavlov, I., Emelyanov, A., Markov, S., et al. (2022) · 2022
Closest in time.
Everything at once-multi-modal fusion transformer for video retrieval
Shvetsova, N., Chen, B., Rouditchenko, A., Thomas, S., Kingsbury, B., Feris, R. S., Harwath, D., Glass, J., and Kuehne, H. (2022) · 2022
Closest in time.
Make-a-video: Text-to-video generation without text-video data
Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., et al. (2022) · 2022
Closest in time.
Flava: A foundational language and vision alignment model
Singh, A., Hu, R., Goswami, V., Couairon, G., Galuba, W., Rohrbach, M., and Kiela, D. (2022) · 2022
Closest in time.
Clip models are few-shot learners: Empirical studies on vqa and visual entailment
Song, H., Dong, L., Zhang, W.-N., Liu, T., and Wei, F. (2022) · 2022
Closest in time.
Worst of both worlds: Biases compound in pre-trained vision-and-language models
Srinivasan, T. and Bisk, Y. (2022) · 2022
Closest in time.
Language models can see: Plugging visual controls in text generation
Su, Y., Lan, T., Liu, Y., Liu, F., Yogatama, D., Wang, Y., Kong, L., and Collier, N. (2022) · 2022
Closest in time.
Zerocap: Zero-shot image-to-text generation for visual-semantic arithmetic
Tewel, Y., Shalev, Y., Schwartz, I., and Wolf, L. (2022) · 2022
Closest in time.
Winoground: Probing vision and language models for visio-linguistic compositionality
Thrush, T., Jiang, R., Bartolo, M., Singh, A., Williams, A., Kiela, D., and Ross, C. (2022) · 2022
Closest in time.
Phenaki: Variable length video generation from open domain textual description
Villegas, R., Babaeizadeh, M., Kindermans, P.-J., Moraldo, H., Zhang, H., Saffar, M. T., Castro, S., Kunze, J., and Erhan, D. (2022) · 2022
Closest in time.
Robust fine-tuning of zero-shot models
Wortsman, M., Ilharco, G., Kim, J. W., Li, M., Kornblith, S., Roelofs, R., Lopes, R. G., Hajishirzi, H., Farhadi, A., Namkoong, H., et al. (2022) · 2022
Closest in time.
Visual clues: Bridging vision and language foundations for image paragraph captioning
Xie, Y., Zhou, L., Dai, X., Yuan, L., Bach, N., Liu, C., and Zeng, M. (2022) · 2022
Closest in time.
Groupvit: Semantic segmentation emerges from text supervision
Xu, J., De Mello, S., Liu, S., Byeon, W., Breuel, T., Kautz, J., and Wang, X. (2022) · 2022
Closest in time.
Advancing high-resolution video-language representation with large-scale video transcriptions
Xue, H., Hang, T., Zeng, Y., Sun, Y., Liu, B., Yang, H., Fu, J., and Guo, B. (2022) · 2022
Closest in time.
Filip: Fine-grained interactive language-image pre-training
Yao, L., Huang, R., Hou, L., Lu, G., Niu, M., Xu, H., Liang, X., Li, Z., Jiang, X., and Xu, C. (2022) · 2022
Closest in time.
A survey of knowledge-intensive nlp with pre-trained language models
Yin, D., Dong, L., Cheng, H., Liu, X., Chang, K.-W., Wei, F., and Gao, J. (2022) · 2022
Closest in time.
Learning visual representation from modality-shared contrastive language-image pre-training
You, H., Zhou, L., Xiao, B., Codella, N., Cheng, Y., Xu, R., Chang, S.-F., and Yuan, L. (2022) · 2022
Closest in time.
Rlip: Relational language-image pre-training for human-object interaction detection
Yuan, H., Jiang, J., Albanie, S., Feng, T., Huang, Z., Ni, D., and Tang, M. (2022) · 2022
Closest in time.
Open-vocabulary detr with conditional matching
Zang, Y., Li, W., Zhou, K., Huang, C., and Loy, C. C. (2022) · 2022
Closest in time.
Merlot reserve: Neural script knowledge through vision and language and sound
Zellers, R., Lu, J., Lu, X., Yu, Y., Zhao, Y., Salehi, M., Kusupati, A., Hessel, J., Farhadi, A., and Choi, Y. (2022) · 2022
Closest in time.
Lit: Zero-shot transfer with locked-image text tuning
Zhai, X., Wang, X., Mustafa, B., Steiner, A., Keysers, D., Kolesnikov, A., and Beyer, L. (2022) · 2022
Closest in time.
P3iv: Probabilistic procedure planning from instructional videos with weak supervision
Zhao, H., Hadji, I., Dvornik, N., Derpanis, K. G., Wildes, R. P., and Jepson, A. D. (2022) · 2022
Closest in time.
Regionclip: Region-based language-image pretraining
Zhong, Y., Yang, J., Zhang, P., Li, C., Codella, N., Li, L. H., Zhou, L., Dai, X., Yuan, L., Li, Y., et al. (2022) · 2022
Closest in time.
Uni-perceiver: Pre-training unified architecture for generic perception for zero-shot and few-shot tasks
Zhu, X., Zhu, J., Li, H., Wu, X., Li, H., Wang, X., and Dai, J. (2022) · 2022
Closest in time.