Fetching the paper…
Reading the bibliography…
This paper presents a comprehensive survey of the taxonomy and evolution of multimodal foundation models that demonstrate vision and vision-language capabilities, focusing on the transition from specialist models to general-purpose assistants.
A theoretical analysis of contrastive unsupervised representation learning
Arora, S., Khandeparkar, H., Khodak, M., Plevrakis, O., and Saunshi, N. (2019) · 1902
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y. (2019) · 1904
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019) · 1907
Earlier work this paper cites.
VisualBERT: A simple and performant baseline for vision and language
Li, L. H., Yatskar, M., Yin, D., Hsieh, C.-J., and Chang, K.-W. (2019b) · 1908
Earlier work this paper cites.
Image co-segmentation via saliency co-fusion
Jerripothula, K. R., Cai, J., and Yuan, J. (2016) · 1909
Earlier work this paper cites.
Knowledge enhanced contextual word representations
Peters, M. E., Neumann, M., Logan IV, R. L., Schwartz, R., Joshi, V., Singh, S., and Smith, N. A. (2019) · 1909
Earlier work this paper cites.
Multi-concept customization of text-to-image diffusion
Kumari, N., Zhang, B., Zhang, R., Shechtman, E., and Zhu, J.-Y. (2023) · 1941
Earlier work this paper cites.
Discriminative clustering for image co-segmentation
Joulin, A., Bach, F., and Ponce, J. (2010) · 1950
Earlier work this paper cites.
Interactive segmentation with intelligent scissors
Mortensen, E. N. and Barrett, W. A. (1998) · 1998
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. (2020) · 2001
Earlier work this paper cites.
Rapid object detection using a boosted cascade of simple features
Viola, P. and Jones, M. (2001) · 2001
Earlier work this paper cites.
Realm: Retrieval-augmented language model pre-training
Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M.-W. (2020) · 2002
Earlier work this paper cites.
Improved baselines with momentum contrastive learning
Chen, X., Fan, H., Girshick, R., and He, K. (2020c) · 2003
Earlier work this paper cites.
Pixel-BERT: Aligning image pixels with text by deep multi-modal transformers
Huang, Z., Zeng, Z., Liu, B., Fu, D., and Fu, J. (2020) · 2004
Earlier work this paper cites.
Object tracking: A survey
Yilmaz, A., Javed, O., and Shah, M. (2006) · 2006
Earlier work this paper cites.
Image retrieval: Ideas, influences, and trends of the new age
Datta, R., Joshi, D., Li, J., and Wang, J. Z. (2008) · 2008
Earlier work this paper cites.
Multi-task learning with deep neural networks: A survey
Crawshaw, M. (2020) · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009) · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M. and Hyvärinen, A. (2010) · 2010
Earlier work this paper cites.
Single image haze removal using dark channel prior
He, K., Sun, J., and Tang, X. (2010) · 2010
Earlier work this paper cites.
A comparative evaluation of interactive segmentation algorithms
McGuinness, K. and O’connor, N. E. (2010) · 2010
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Shin, T., Razeghi, Y., Logan IV, R. L., Wallace, E., and Singh, S. (2020) · 2010
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Ordonez, V., Kulkarni, G., and Berg, T. (2011) · 2011
Earlier work this paper cites.
The pascal visual object classes challenge 2012 (voc2012) development kit
Everingham, M. and Winn, J. (2011) · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012) · 2012
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
Silberman, N., Hoiem, D., Kohli, P., and Fergus, R. (2012) · 2012
Earlier work this paper cites.
Semantic parsing on freebase from question-answer pairs
Berant, J., Chou, A., Frostig, R., and Liang, P. (2013) · 2013
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Frome, A., Corrado, G. S., Shlens, J., Bengio, S., Dean, J., Ranzato, M., and Mikolov, T. (2013) · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M. (2013) · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Mikolov, T., Chen, K., Corrado, G., and Dean, J. (2013) · 2013
Earlier work this paper cites.
Online object tracking: A benchmark
Wu, Y., Lim, J., and Yang, M.-H. (2013) · 2013
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
Kazemzadeh, S., Ordonez, V., Matten, M., and Berg, T. (2014) · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L. (2014) · 2014
Earlier work this paper cites.
The role of context for object detection and semantic segmentation in the wild
Mottaghi, R., Chen, X., Liu, X., Cho, N.-G., Lee, S.-W., Fidler, S., Urtasun, R., and Yuille, A. (2014) · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D. (2015) · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y. (2015) · 2015
Earlier work this paper cites.
Microsoft COCO captions: Data collection and evaluation server
Chen, X., Fang, H., Lin, T., Vedantam, R., Gupta, S., Dollár, P., and Zitnick, C. L. (2015) · 2015
Earlier work this paper cites.
Fast r-cnn
Girshick, R. (2015) · 2015
Earlier work this paper cites.
Region-based convolutional networks for accurate object detection and segmentation
Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2015) · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Long, J., Shelhamer, E., and Darrell, T. (2015) · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J., and Lazebnik, S. (2015) · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, S., He, K., Girshick, R., and Sun, J. (2015) · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T. (2015) · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al. (2015) · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S. (2015) · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Vinyals, O., Toshev, A., Bengio, S., and Erhan, D. (2015) · 2015
Earlier work this paper cites.
Show and tell: Lessons learned from the 2015 mscoco image captioning challenge
Vinyals, O., Toshev, A., Bengio, S., and Erhan, D. (2016) · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Earlier work this paper cites.
Segmentation from natural language expressions
Hu, R., Rohrbach, M., and Darrell, T. (2016) · 2016
Earlier work this paper cites.
Discriminative regularization for generative models
Lamb, A., Dumoulin, V., and Courville, A. (2016) · 2016
Earlier work this paper cites.
Autoencoding beyond pixels using a learned similarity metric
Larsen, A. B. L., Sønderby, S. K., Larochelle, H., and Winther, O. (2016) · 2016
Earlier work this paper cites.
Ssd: Single shot multibox detector
Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.-Y., and Berg, A. C. (2016) · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Mao, J., Huang, J., Toshev, A., Camburu, O., Yuille, A. L., and Murphy, K. (2016) · 2016
Earlier work this paper cites.
Cross-stitch networks for multi-task learning
Misra, I., Shrivastava, A., Gupta, A., and Hebert, M. (2016) · 2016
Earlier work this paper cites.
Modeling context between objects for referring expression understanding
Nagaraja, V. K., Morariu, V. I., and Davis, L. S. (2016) · 2016
Earlier work this paper cites.
You only look once: Unified, real-time object detection
Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016) · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A. (2016) · 2016
Earlier work this paper cites.
Yfcc100m: The new data in multimedia research
Thomee, B., Shamma, D. A., Friedland, G., Elizalde, B., Ni, K., Poland, D., Borth, D., and Li, L.-J. (2016) · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Yu, L., Poirson, P., Yang, S., Berg, A. C., and Berg, T. L. (2016) · 2016
Earlier work this paper cites.
Text summarization techniques: a brief survey
Allahyari, M., Pouriyeh, S., Assefi, M., Safaei, S., Trippe, E. D., Gutierrez, J. B., and Kochut, K. (2017) · 2017
Earlier work this paper cites.
Annotating object instances with a polygon-rnn
Castrejon, L., Kundu, K., Urtasun, R., and Fidler, S. (2017) · 2017
Earlier work this paper cites.
Rethinking atrous convolution for semantic image segmentation
Chen, L.-C., Papandreou, G., Schroff, F., and Adam, H. (2017) · 2017
Earlier work this paper cites.
Mask r-cnn
He, K., Gkioxari, G., Dollár, P., and Girshick, R. (2017) · 2017
Earlier work this paper cites.
Ubernet: Training a universal convolutional neural network for low-, mid-, and high-level vision using diverse datasets and limited memory
Kokkinos, I. (2017) · 2017
Earlier work this paper cites.
Focal loss for dense object detection
Lin, T.-Y., Goyal, P., Girshick, R., He, K., and Dollár, P. (2017) · 2017
Earlier work this paper cites.
Recurrent multimodal interaction for referring image segmentation
Liu, C., Lin, Z., Shen, X., Yang, J., Lu, X., and Yuille, A. (2017) · 2017
Earlier work this paper cites.
Neural discrete representation learning
Oord, A. v. d., Vinyals, O., and Kavukcuoglu, K. (2017) · 2017
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era
Sun, C., Shrivastava, A., Singh, S., and Gupta, A. (2017) · 2017
Earlier work this paper cites.
Neural discrete representation learning
van den Oord, A., Vinyals, O., and Kavukcuoglu, K. (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Scene parsing through ade20k dataset
Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., and Torralba, A. (2017) · 2017
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., and Zhang, L. (2018) · 2018
Earlier work this paper cites.
Zero-shot object detection
Bansal, A., Sikka, K., Sharma, G., Chellappa, R., and Divakaran, A. (2018) · 2018
Earlier work this paper cites.
Deep clustering for unsupervised learning of visual features
Caron, M., Bojanowski, P., Joulin, A., and Douze, M. (2018) · 2018
Earlier work this paper cites.
Generative adversarial networks: An overview
Creswell, A., White, T., Dumoulin, V., Arulkumaran, K., Sengupta, B., and Bharath, A. A. (2018) · 2018
Earlier work this paper cites.
Visual grounding via accumulated attention
Deng, C., Wu, Q., Wu, Q., Hu, F., Lyu, F., and Tan, M. (2018) · 2018
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
Gurari, D., Li, Q., Stangl, A. J., Guo, A., Lin, C., Grauman, K., Luo, J., and Bigham, J. P. (2018) · 2018
Earlier work this paper cites.
Learning deep representations by mutual information estimation and maximization
Hjelm, R. D., Fedorov, A., Lavoie-Marchildon, S., Grewal, K., Bachman, P., Trischler, A., and Bengio, Y. (2018) · 2018
Earlier work this paper cites.
Dynamic multimodal instance segmentation guided by natural language queries
Margffoy-Tuay, E., Pérez, J. C., Botero, E., and Arbeláez, P. (2018) · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O. (2018) · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R. (2018) · 2018
Earlier work this paper cites.
Cosface: Large margin cosine loss for deep face recognition
Wang, H., Wang, Y., Zhou, Z., Ji, X., Li, Z., Gong, D., Zhou, J., and Liu, W. (2018) · 2018
Earlier work this paper cites.
Unsupervised feature learning via non-parametric instance discrimination
Wu, Z., Xiong, Y., Yu, S. X., and Lin, D. (2018) · 2018
Earlier work this paper cites.
Zero-shot learning—a comprehensive evaluation of the good, the bad and the ugly
Xian, Y., Lampert, C. H., Schiele, B., and Akata, Z. (2018) · 2018
Earlier work this paper cites.
Pad-net: Multi-tasks guided prediction-and-distillation network for simultaneous depth estimation and scene parsing
Xu, D., Ouyang, W., Wang, X., and Sebe, N. (2018) · 2018
Earlier work this paper cites.
Taskonomy: Disentangling task transfer learning
Zamir, A. R., Sax, A., Shen, W., Guibas, L. J., Malik, J., and Savarese, S. (2018) · 2018
Earlier work this paper cites.
Generative domain-migration hashing for sketch-to-image retrieval
Zhang, J., Shen, F., Liu, L., Zhu, F., Yu, M., Shao, L., Shen, H. T., and Van Gool, L. (2018) · 2018
Earlier work this paper cites.
nocaps: novel object captioning at scale
Agrawal, H., Desai, K., Wang, Y., Chen, X., Jain, R., Johnson, M., Batra, D., Parikh, D., Lee, S., and Anderson, P. (2019) · 2019
Earlier work this paper cites.
Learning representations by maximizing mutual information across views
Bachman, P., Hjelm, R. D., and Buchwalter, W. (2019) · 2019
Earlier work this paper cites.
Yolact: Real-time instance segmentation
Bolya, D., Zhou, C., Xiao, F., and Lee, Y. J. (2019) · 2019
Earlier work this paper cites.
Object grounding via iterative context reasoning
Chen, L., Zhai, M., He, J., and Mori, G. (2019) · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019) · 2019
Earlier work this paper cites.
Panoptic segmentation
Kirillov, A., He, K., Girshick, R., Rother, C., and Dollár, P. (2019) · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J., Batra, D., Parikh, D., and Lee, S. (2019) · 2019
Earlier work this paper cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Miech, A., Zhukov, D., Alayrac, J.-B., Tapaswi, M., Laptev, I., and Sivic, J. (2019) · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. (2019) · 2019
Earlier work this paper cites.
Generating diverse high-fidelity images with vq-vae-2
Razavi, A., Van den Oord, A., and Vinyals, O. (2019) · 2019
Earlier work this paper cites.
Objects365: A large-scale, high-quality dataset for object detection
Shao, S., Li, Z., Zhang, T., Peng, C., Yu, G., Zhang, X., Li, J., and Sun, J. (2019) · 2019
Earlier work this paper cites.
VL-BERT: Pre-training of generic visual-linguistic representations
Su, W., Zhu, X., Cao, Y., Li, B., Lu, L., Wei, F., and Dai, J. (2019) · 2019
Earlier work this paper cites.
LXMERT: Learning cross-modality encoder representations from transformers
Tan, H. and Bansal, M. (2019) · 2019
Earlier work this paper cites.
Deep modular co-attention networks for visual question answering
Yu, Z., Yu, J., Cui, Y., Tao, D., and Tian, Q. (2019) · 2019
Earlier work this paper cites.
From recognition to cognition: Visual commonsense reasoning
Zellers, R., Bisk, Y., Farhadi, A., and Choi, Y. (2019) · 2019
Earlier work this paper cites.
Zero shot detection
Zhu, P., Wang, H., and Saligrama, V. (2019) · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 2020
Earlier work this paper cites.
End-to-end object detection with transformers
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. (2020) · 2020
Earlier work this paper cites.
Unsupervised learning of visual features by contrasting cluster assignments
Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., and Joulin, A. (2020) · 2020
Earlier work this paper cites.
Large-scale adversarial training for vision-and-language representation learning
Gan, Z., Chen, Y.-C., Li, L., Zhu, C., Cheng, Y., and Liu, J. (2020) · 2020
Earlier work this paper cites.
Generative adversarial networks
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2020) · 2020
Earlier work this paper cites.
Bootstrap your own latent-a new approach to self-supervised learning
Grill, J.-B., Strub, F., Altché, F., Tallec, C., Richemond, P., Buchatskaya, E., Doersch, C., Avila Pires, B., Guo, Z., Gheshlaghi Azar, M., et al. (2020) · 2020
Earlier work this paper cites.
A survey on instance segmentation: state of the art
Hafiz, A. M. and Bhat, G. M. (2020) · 2020
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. (2020) · 2020
Earlier work this paper cites.
Data-efficient image recognition with contrastive predictive coding
Henaff, O. (2020) · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P. (2020) · 2020
Cited alongside, same era.
A survey on contrastive self-supervised learning
Jaiswal, A., Babu, A. R., Zadeh, M. Z., Banerjee, D., and Makedon, F. (2020) · 2020
Cited alongside, same era.
Self-supervised visual feature learning with deep neural networks: A survey
Jing, L. and Tian, Y. (2020) · 2020
Cited alongside, same era.
Big transfer (bit): General visual representation learning
Kolesnikov, A., Beyer, L., Zhai, X., Puigcerver, J., Yung, J., Gelly, S., and Houlsby, N. (2020) · 2020
Cited alongside, same era.
Mseg: A composite dataset for multi-domain semantic segmentation
Lambert, J., Liu, Z., Sener, O., Hays, J., and Koltun, V. (2020) · 2020
Multimodal open-vocabulary video classification via pre-trained vision and language models
Qian, R., Li, Y., Xu, Z., Yang, M.-H., Belongie, S., and Cui, Y. (2022) · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. (2022) · 2022
Later among the works it cites.
Denseclip: Language-guided dense prediction with context-aware prompting
Rao, Y., Zhao, W., Chen, G., Tang, Y., Zhu, Z., Huang, G., Zhou, J., and Lu, J. (2022) · 2022
Later among the works it cites.
Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., et al. (2022) · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al. (2020) · 2020
Cited alongside, same era.
12-in-1: Multi-task vision and language representation learning
Lu, J., Goswami, V., Rohrbach, M., Parikh, D., and Lee, S. (2020) · 2020
Cited alongside, same era.
Self-supervised learning of pretext-invariant representations
Misra, I. and Maaten, L. v. d. (2020) · 2020
Cited alongside, same era.
A metric learning reality check
Musgrave, K., Belongie, S., and Lim, S.-N. (2020) · 2020
Cited alongside, same era.
Connecting vision and language with localized narratives
Pont-Tuset, J., Uijlings, J., Changpinyo, S., Soricut, R., and Ferrari, V. (2020) · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. (2020) · 2020
Cited alongside, same era.
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. (2022) · 2022
Later among the works it cites.
Survey: Transformer based video-language pre-training
Ruan, L. and Jin, Q. (2022) · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., et al. (2022) · 2022
Later among the works it cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al. (2022) · 2022
Later among the works it cites.
A-okvqa: A benchmark for visual question answering using world knowledge
Schwenk, D., Khandelwal, A., Clark, C., Marino, K., and Mottaghi, R. (2022) · 2022
Later among the works it cites.
Knn-diffusion: Image generation via large-scale retrieval
Sheynin, S., Ashual, O., Polyak, A., Singer, U., Gafni, O., Nachmani, E., and Taigman, Y. (2022) · 2022
Later among the works it cites.
Make-a-video: Text-to-video generation without text-video data
Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., et al. (2022) · 2022
Later among the works it cites.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Tong, Z., Song, Y., Wang, J., and Wang, L. (2022) · 2022
Later among the works it cites.
Retrieval-augmented multimodal language modeling
Yasunaga, M., Aghajanyan, A., Shi, W., James, R., Leskovec, J., Liang, P., Lewis, M., Zettlemoyer, L., and Yih, W.-t. (2022) · 2022
Later among the works it cites.
Masked image modeling with denoising contrast
Yi, K., Ge, Y., Li, X., Yang, S., Li, D., Wu, J., Shan, Y., and Qie, X. (2022) · 2022
Later among the works it cites.
A survey of knowledge-intensive nlp with pre-trained language models
Yin, D., Dong, L., Cheng, H., Liu, X., Chang, K.-W., Wei, F., and Gao, J. (2022) · 2022
Later among the works it cites.
Open-vocabulary detr with conditional matching
Zang, Y., Li, W., Zhou, K., Huang, C., and Loy, C. C. (2022) · 2022
Later among the works it cites.
Multi-grained vision language pre-training: Aligning texts with visual concepts
Zeng, Y., Zhang, X., and Li, H. (2022) · 2022
Later among the works it cites.
End-to-end instance edge detection
Zou, X., Liu, H., and Lee, Y. J. (2022) · 2022
Later among the works it cites.
A-star: Test-time attention segregation and retention for text-to-image synthesis
Agarwal, A., Karanam, S., Joseph, K., Saxena, A., Goswami, K., and Srinivasan, B. V. (2023) · 2023
Closest in time.
Openflamingo
Awadalla, A., Gao, I., Gardner, J., Hessel, J., Hanafy, Y., Zhu, W., Marathe, K., Bitton, Y., Gadre, S., Jitsev, J., Kornblith, S., Koh, P. W., Ilharco, G., Wortsman, M., and Schmidt, L. (2023) · 2023
Closest in time.
Foundational models defining a new era in vision: A survey and outlook
Awais, M., Naseer, M., Khan, S., Anwer, R. M., Cholakkal, H., Shah, M., Yang, M.-H., and Khan, F. S. (2023) · 2023
Closest in time.
Towards in-context scene understanding
Balažević, I., Steiner, D., Parthasarathy, N., Arandjelović, R., and Hénaff, O. J. (2023) · 2023
Closest in time.
Universal guidance for diffusion models
Bansal, A., Chu, H.-M., Schwarzschild, A., Sengupta, S., Goldblum, M., Geiping, J., and Goldstein, T. (2023) · 2023
Closest in time.
Visit-bench: A benchmark for vision-language instruction following inspired by real-world use
Bitton, Y., Bansal, H., Hessel, J., Shao, R., Zhu, W., Awadalla, A., Gardner, J., Taori, R., and Schimdt, L. (2023) · 2023
Closest in time.
Training diffusion models with reinforcement learning
Black, K., Janner, M., Du, Y., Kostrikov, I., and Levine, S. (2023) · 2023
Closest in time.
Instructpix2pix: Learning to follow image editing instructions
Brooks, T., Holynski, A., and Efros, A. A. (2023) · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al. (2023) · 2023
Closest in time.
Large language models as tool makers
Cai, T., Wang, X., Ma, T., Chen, X., and Zhou, D. (2023) · 2023
Closest in time.
Less is more: Removing text-regions improves clip training efficiency and robustness
Cao, L., Zhang, B., Chen, C., Yang, Y., Du, X., Zhang, W., Lu, Z., and Zheng, Y. (2023) · 2023
Closest in time.
Muse: Text-to-image generation via masked generative transformers
Chang, H., Zhang, H., Barber, J., Maschinot, A., Lezama, J., Jiang, L., Yang, M.-H., Murphy, K., Freeman, W. T., Rubinstein, M., et al. (2023) · 2023
Closest in time.
Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models
Chefer, H., Alaluf, Y., Vinker, Y., Wolf, L., and Cohen-Or, D. (2023) · 2023
Closest in time.
Reproducible scaling laws for contrastive language-image learning
Cherti, M., Beaumont, R., Wightman, R., Wortsman, M., Ilharco, G., Gordon, C., Schuhmann, C., Schmidt, L., and Jitsev, J. (2023) · 2023
Closest in time.
Diagnostic benchmark and iterative inpainting for layout-guided image generation
Cho, J., Li, L., Yang, Z., Gan, Z., Wang, L., and Bansal, M. (2023) · 2023
Closest in time.
Redpajama-data: An open source recipe to reproduce llama training dataset
Computer, T. (2023) · 2023
Closest in time.
Qlora: Efficient finetuning of quantized llms
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. (2023) · 2023
Closest in time.
Peco: Perceptual codebook for bert pre-training of vision transformers
Dong, X., Bao, J., Zhang, T., Chen, D., Zhang, W., Yuan, L., Chen, D., Wen, F., Yu, N., and Guo, B. (2023) · 2023
Closest in time.
PaLM-E: An embodied multimodal language model
Driess, D., Xia, F., Sajjadi, M. S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al. (2023) · 2023
Closest in time.
Eva: Exploring the limits of masked visual representation learning at scale
Fang, Y., Wang, W., Xie, B., Sun, Q., Wu, L., Wang, X., Huang, T., Wang, X., and Cao, Y. (2023) · 2023
Closest in time.
Layoutgpt: Compositional visual planning and generation with large language models
Feng, W., Zhu, W., Fu, T.-j., Jampani, V., Akula, A., He, X., Basu, S., Wang, X. E., and Wang, W. Y. (2023) · 2023
Closest in time.
Mme: A comprehensive evaluation benchmark for multimodal large language models
Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., Qiu, Z., Lin, W., Yang, J., Zheng, X., et al. (2023) · 2023
Closest in time.
Datacomp: In search of the next generation of multimodal datasets
Gadre, S. Y., Ilharco, G., Fang, A., Hayase, J., Smyrnis, G., Nguyen, T., Marten, R., Wortsman, M., Ghosh, D., Zhang, J., et al. (2023) · 2023
Closest in time.
Planting a seed of vision in large language model
Ge, Y., Ge, Y., Zeng, Z., Wang, X., and Shan, Y. (2023) · 2023
Closest in time.
Openllama: An open reproduction of llama
Geng, X. and Liu, H. (2023) · 2023
Closest in time.
Instructdiffusion: A generalist modeling interface for vision tasks
Geng, Z., Yang, B., Hang, T., Li, C., Gu, S., Zhang, T., Bao, J., Zhang, Z., Hu, H., Chen, D., et al. (2023) · 2023
Closest in time.
Imagebind: One embedding space to bind them all
Girdhar, R., El-Nouby, A., Liu, Z., Singh, M., Alwala, K. V., Joulin, A., and Misra, I. (2023) · 2023
Closest in time.
Multimodal-gpt: A vision and language model for dialogue with humans
Gong, T., Lyu, C., Zhang, S., Wang, Y., Zheng, M., Zhao, Q., Liu, K., Zhang, W., Luo, P., and Chen, K. (2023) · 2023
Closest in time.
Dataseg: Taming a universal multi-dataset multi-task segmentation model
Gu, X., Cui, Y., Huang, J., Rashwan, A., Yang, X., Zhou, X., Ghiasi, G., Kuo, W., Chen, H., Chen, L.-C., et al. (2023) · 2023
Closest in time.
The false promise of imitating proprietary llms
Gudibande, A., Wallace, E., Snell, C., Geng, X., Liu, H., Abbeel, P., Levine, S., and Song, D. (2023) · 2023
Closest in time.
Detecting and preventing hallucinations in large vision language models
Gunjal, A., Yin, J., and Bas, E. (2023) · 2023
Closest in time.
Visual programming: Compositional visual reasoning without training
Gupta, T. and Kembhavi, A. (2023) · 2023
Closest in time.
3d-llm: Injecting the 3d world into large language models
Hong, Y., Zhen, H., Chen, P., Zheng, S., Du, Y., Chen, Z., and Gan, C. (2023) · 2023
Closest in time.
Bliva: A simple multimodal llm for better handling of text-rich visual questions
Hu, W., Xu, Y., Li, Y., Li, W., Chen, Z., and Tu, Z. (2023) · 2023
Closest in time.
Oneformer: One transformer to rule universal image segmentation
Jain, J., Li, J., Chiu, M. T., Hassani, A., Orlov, N., and Shi, H. (2023) · 2023
Closest in time.
Scaling up gans for text-to-image synthesis
Kang, M., Zhu, J.-Y., Zhang, R., Park, J., Shechtman, E., Paris, S., and Park, T. (2023) · 2023
Closest in time.
Imagic: Text-based real image editing with diffusion models
Kawar, B., Zada, S., Lang, O., Tov, O., Chang, H., Dekel, T., Mosseri, I., and Irani, M. (2023) · 2023
Closest in time.
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., et al. (2023) · 2023
Closest in time.
Generating images with multimodal language models
Koh, J. Y., Fried, D., and Salakhutdinov, R. (2023) · 2023
Closest in time.
Lisa: Reasoning segmentation via large language model
Lai, X., Tian, Z., Chen, Y., Li, Y., Yuan, Y., Liu, S., and Jia, J. (2023) · 2023
Closest in time.
Obelisc: An open web-scale filtered dataset of interleaved image-text documents
Laurençon, H., Saulnier, L., Tronchon, L., Bekman, S., Singh, A., Lozhkov, A., Wang, T., Karamcheti, S., Rush, A. M., Kiela, D., et al. (2023) · 2023
Closest in time.
Scigraphqa: A large-scale synthetic multi-turn question-answering dataset for scientific graphs
Li, S. and Tajbakhsh, N. (2023) · 2023
Closest in time.
Lin, Y., Wu, H., Wang, R., Lu, H., Lin, X., Xiong, H., and Wang, L. (2023) · 2023
Closest in time.
Stone needle: A general multimodal large-scale model framework towards healthcare
Liu, W. and Zuo, Y. (2023) · 2023
Closest in time.
Segment anything in medical images
Ma, J. and Wang, B. (2023) · 2023
Closest in time.
Can sam count anything? an empirical study on sam counting
Ma, Z., Hong, X., and Shangguan, Q. (2023) · 2023
Closest in time.
Metavl: Transferring in-context learning ability from language models to vision-language models
Monajatipoor, M., Li, L. H., Rouhsedaghat, M., Yang, L. F., and Chang, K.-W. (2023) · 2023
Closest in time.
Med-flamingo: a multimodal medical few-shot learner
Moor, M., Huang, Q., Wu, S., Yasunaga, M., Zakka, C., Dalmia, Y., Reis, E. P., Rajpurkar, P., and Leskovec, J. (2023) · 2023
Closest in time.
Mou, C., Wang, X., Xie, L., Zhang, J., Qi, Z., Shan, Y., and Qie, X. (2023) · 2023
Closest in time.
Embodiedgpt: Vision-language pre-training via embodied chain of thought
Mu, Y., Zhang, Q., Hu, M., Wang, W., Ding, M., Jin, J., Wang, B., Dai, J., Qiao, Y., and Luo, P. (2023) · 2023
Closest in time.
All in tokens: Unifying output space of visual tasks via soft token
Ning, J., Li, C., Zhang, Z., Geng, Z., Dai, Q., He, K., and Hu, H. (2023) · 2023
Closest in time.
Dinov2: Learning robust visual features without supervision
Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al. (2023) · 2023
Closest in time.
Know your self-supervised learning: A survey on image-based generative and discriminative training
Ozbulak, U., Lee, H. J., Boga, B., Anzaku, E. T., Park, H., Van Messem, A., De Neve, W., and Vankerschaver, J. (2023) · 2023
Closest in time.
Art: Automatic multi-step reasoning and tool-use for large language models
Paranjape, B., Lundberg, S., Singh, S., Hajishirzi, H., Zettlemoyer, L., and Ribeiro, M. T. (2023) · 2023
Closest in time.
Gorilla: Large language model connected with massive apis
Patil, S. G., Zhang, T., Wang, X., and Gonzalez, J. E. (2023) · 2023
Closest in time.
Penedo, G., Malartic, Q., Hesslow, D., Cojocaru, R., Cappelli, A., Alobeidli, H., Pannier, B., Almazrouei, E., and Launay, J. (2023) · 2023
Closest in time.
Detgpt: Detect what you need via reasoning
Pi, R., Gao, J., Diao, S., Pan, R., Dong, H., Zhang, J., Yao, L., Han, J., Xu, H., and Zhang, L. K. T. (2023) · 2023
Closest in time.
Qian, C., Han, C., Fung, Y. R., Qin, Y., Liu, Z., and Ji, H. (2023) · 2023
Closest in time.
Segment anything meets point tracking
Rajič, F., Ke, L., Tai, Y.-W., Tang, C.-K., Danelljan, M., and Yu, F. (2023) · 2023
Closest in time.
Sam. md: Zero-shot medical image segmentation capabilities of the segment anything model
Roy, S., Wald, T., Koehler, G., Rokuss, M. R., Disch, N., Holzschuh, J., Zimmerer, D., and Maier-Hein, K. H. (2023) · 2023
Closest in time.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., and Aberman, K. (2023) · 2023
Closest in time.
Toolformer: Language models can teach themselves to use tools
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., and Scialom, T. (2023) · 2023
Closest in time.
Tiny lvlm-ehub: Early multimodal experiments with bard
Shao, W., Hu, Y., Gao, P., Lei, M., Zhang, K., Meng, F., Xu, P., Huang, S., Li, H., Qiao, Y., et al. (2023) · 2023
Closest in time.
https://sharegpt.com/
ShareGPT (2023) · 2023
Closest in time.
The effectiveness of mae pre-pretraining for billion-scale pretraining
Singh, M., Duval, Q., Alwala, K. V., Fan, H., Aggarwal, V., Adcock, A., Joulin, A., Dollár, P., Feichtenhofer, C., Girshick, R., et al. (2023) · 2023
Closest in time.
Restgpt: Connecting large language models with real-world applications via restful apis
Song, Y., Xiong, W., Zhu, D., Li, C., Wang, K., Tian, Y., and Li, S. (2023) · 2023
Closest in time.
Pandagpt: One model to instruction-follow them all
Su, Y., Lan, T., Li, H., Xu, J., Wang, Y., and Cai, D. (2023) · 2023
Closest in time.
Vipergpt: Visual inference via python execution for reasoning
Surís, D., Menon, S., and Vondrick, C. (2023) · 2023
Closest in time.
Siamese image modeling for self-supervised vision representation learning
Tao, C., Zhu, X., Su, W., Huang, G., Li, B., Zhou, J., Qiao, Y., Wang, X., and Dai, J. (2023) · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B. (2023) · 2023
Closest in time.
Introducing mpt-7b: A new standard for open-source, ly usable llms
Team, M. N. (2023) · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al. (2023) · 2023
Closest in time.
Image captioners are scalable vision learners too
Tschannen, M., Kumar, M., Steiner, A., Zhai, X., Houlsby, N., and Beyer, L. (2023) · 2023
Closest in time.
Towards generalist biomedical ai
Tu, T., Azizi, S., Driess, D., Schaekermann, M., Amin, M., Chang, P.-C., Carroll, A., Lau, C., Tanno, R., Ktena, I., et al. (2023) · 2023
Closest in time.
Vicuna: An open-source chatbot impressing GPT-4 with 90%* chatgpt quality
Vicuna (2023) · 2023
Closest in time.
Masked autoencoding does not help natural language supervision at scale
Weers, F., Shankar, V., Katharopoulos, A., Yang, Y., and Gunter, T. (2023) · 2023
Closest in time.
Elite: Encoding visual concepts into textual embeddings for customized text-to-image generation
Wei, Y., Zhang, Y., Ji, Z., Bai, J., Zhang, L., and Zuo, W. (2023) · 2023
Closest in time.
Llm-powered autonomous agents
Weng, L. (2023) · 2023
Closest in time.
Instruction-vit: Multi-modal prompts for instruction learning in vit
Xiao, Z., Chen, Y., Zhang, L., Yao, J., Wu, Z., Yu, X., Pan, Y., Zhao, L., Ma, C., Liu, X., et al. (2023) · 2023
Closest in time.
Universal instance perception as object discovery and retrieval
Yan, B., Jiang, Y., Wu, J., Wang, D., Luo, P., Yuan, Z., and Lu, H. (2023) · 2023
Closest in time.
Mm-react: Prompting chatgpt for multimodal reasoning and action
Yang*, Z., Li*, L., Wang*, J., Lin*, K., Azarnasab*, E., Ahmed*, F., Liu, Z., Liu, C., Zeng, M., and Wang, L. (2023) · 2023
Closest in time.
Detclipv2: Scalable open-vocabulary object detection pre-training via word-region alignment
Yao, L., Han, J., Liang, X., Xu, D., Zhang, W., Li, Z., and Xu, H. (2023) · 2023
Closest in time.
Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark
Yin, Z., Wang, J., Cao, J., Shi, Z., Liu, D., Li, M., Sheng, L., Bai, L., Huang, X., Wang, Z., et al. (2023) · 2023
Closest in time.
Scaling autoregressive multi-modal models: Pretraining and instruction tuning
Yu, L. and et al (2023) · 2023
Closest in time.
Contextual object detection with multimodal large language models
Zang, Y., Li, W., Han, J., Zhou, K., and Loy, C. C. (2023) · 2023
Closest in time.
Scenecomposer: Any-level semantic image synthesis
Zeng, Y., Lin, Z., Zhang, J., Liu, Q., Collomosse, J., Kuen, J., and Patel, V. M. (2023) · 2023
Closest in time.
Sigmoid loss for language image pre-training
Zhai, X., Mustafa, B., Kolesnikov, A., and Beyer, L. (2023) · 2023
Closest in time.
Adding conditional control to text-to-image diffusion models
Zhang, L. and Agrawala, M. (2023) · 2023
Closest in time.
How segment anything model (sam) boost medical image segmentation?
Zhang, Y. and Jiao, R. (2023) · 2023
Closest in time.
Vision + language applications: A survey
Zhou, Y. and Shimada, N. (2023) · 2023
Closest in time.
Detrs with collaborative hybrid assignments training
Zong, Z., Song, G., and Liu, Y. (2023) · 2023
Closest in time.
Image inpainting: A review
Elharrouss, O., Almaadeed, N., Al-Maadeed, S., and Akbari, Y. (2020) · 2028
Closest in time.