J. Xu, Y. Ren, H. Tang, Z. Yang, L. Pan, Y. Yang, X. Pu, S. Y. Philip, and L. He, “Self-supervised discriminative feature learning for deep multi-view clustering,” TKDE , 2022
2022
Later among the works it cites.
A. Guzhov, F. Raue, J. Hees, and A. R. Dengel, “Audioclip: Extending clip to image, text and audio,” ICASSP , 2022
2022
Later among the works it cites.
N. Mu, A. Kirillov, D. Wagner, and S. Xie, “Slip: Self-supervision meets language-image pre-training,” in ECCV , 2022
2022
Later among the works it cites.
M. Afham, I. Dissanayake, D. Dissanayake, A. Dharmasiri, K. Thilakarathna, and R. Rodrigo, “Crosspoint: Self-supervised cross-modal contrastive learning for 3d point cloud understanding,” CVPR , 2022
2022
Later among the works it cites.
J. Zhou, J. Wang, J. Zhang, W. Sun, J. Zhang, S. Birchfield, D. Guo, L. Kong, M. Wang, and Y. Zhong, “Audio–visual segmentation,” in ECCV , 2022
2022
Later among the works it cites.
J. Li, D. Li, C. Xiong, and S. C. H. Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” in ICML , 2022
2022
Later among the works it cites.
W.-N. Hsu and B. Shi, “u-hubert: Unified mixed-modal speech pretraining and zero-shot transfer to unlabeled modality,” NeurIPS , 2022
2022
Later among the works it cites.
H. Bao, W. Wang, L. Dong, and F. Wei, “Vl-beit: Generative vision-language pretraining,” arXiv , 2022
2022
Later among the works it cites.
Z. Wang, J. Yu, A. W. Yu, Z. Dai, Y. Tsvetkov, and Y. Cao, “Simvlm: Simple visual language model pretraining with weak supervision,” ICLR , 2022
2022
Later among the works it cites.
W. Wang, H. Bao, L. Dong, and F. Wei, “Vlmo: Unified vision-language pre-training with mixture-of-modality-experts,” NeurIPS , 2022
2022
Later among the works it cites.
R. Zellers, J. Lu, X. Lu, Y. Yu, Y. Zhao, M. Salehi, A. Kusupati, J. Hessel, A. Farhadi, and Y. Choi, “Merlot reserve: Neural script knowledge through vision and language and sound,” CVPR , 2022
2022
Later among the works it cites.
M. Yasunaga, A. Bosselut, H. Ren, X. Zhang, C. D. Manning, P. Liang, and J. Leskovec, “Deep bidirectional language-knowledge graph pretraining,” in NeurIPS , 2022
2022
Later among the works it cites.
Y. Gong, A. H. Liu, A. Rouditchenko, and J. Glass, “Uavm: Towards unifying audio and visual models,” Signal Processing Letters , 2022
2022
Later among the works it cites.
T. Wang, W. Jiang, Z. Lu, F. Zheng, R. Cheng, C. Yin, and P. Luo, “Vlmixer: Unpaired vision-language pre-training via cross-modal cutmix,” in ICML , 2022
2022
Later among the works it cites.
J. Carreira, S. Koppula, D. Zoran, A. Recasens, C. Ionescu, O. Henaff, E. Shelhamer, R. Arandjelovic, M. Botvinick, O. Vinyals et al. , “Hierarchical perceiver,” arXiv , 2022
2022
Later among the works it cites.
Y. Dai, D. Tang, L. Liu, M. Tan, C. Zhou, J. Wang, Z. Feng, F. Zhang, X. Hu, and S. Shi, “One model, multiple modalities: A sparsely activated approach for text, sound, image, video and code,” arXiv , 2022
2022
Later among the works it cites.
N. Shvetsova, B. Chen, A. Rouditchenko, S. Thomas, B. Kingsbury, R. S. Feris, D. Harwath, J. Glass, and H. Kuehne, “Everything at once-multi-modal fusion transformer for video retrieval,” in CVPR , 2022
2022
Later among the works it cites.
P. H. Seo, A. Nagrani, A. Arnab, and C. Schmid, “End-to-end generative pretraining for multimodal video captioning,” in CVPR , 2022
2022
Later among the works it cites.
C. Eichenberg, S. Black, S. Weinbach, L. Parcalabescu, and A. Frank, “MAGMA – multimodal augmentation of generative models through adapter-based finetuning,” in Findings of EMNLP , 2022
2022
Later among the works it cites.
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al. , “Flamingo: a visual language model for few-shot learning,” NeurIPS , 2022
2022
Later among the works it cites.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al. , “Chain-of-thought prompting elicits reasoning in large language models,” NeurIPS , 2022
2022
Later among the works it cites.
M. Zhou, L. Yu, A. Singh, M. M. Wang, Z. Yu, and N. Zhang, “Unsupervised vision-and-language pretraining via retrieval-based multi-granular alignment,” CVPR , 2022
2022
Later among the works it cites.
Y. Huang, Y. Wang, Y. Zeng, and L. Wang, “Mack: Multimodal aligned conceptual knowledge for unpaired image-text matching,” in NeurIPS , 2022
2022
Later among the works it cites.
Q. shi Zhu, L. Zhou, Z.-H. Zhang, S. Liu, B. Jiao, J. Zhang, L. Dai, D. Jiang, J. Li, and F. Wei, “Vatlm: Visual-audio-text pre-training with unified masked prediction for speech representation learning,” Transactions on Multimedia , 2022
2022
Later among the works it cites.
C. Sautier, G. Puy, S. Gidaris, A. Boulch, A. Bursuc, and R. Marlet, “Image-to-lidar self-supervised distillation for autonomous driving data,” CVPR , 2022
2022
Later among the works it cites.
N. Botteghi, M. Poel, and C. Brune, “Unsupervised representation learning in deep reinforcement learning: A review,” arXiv , 2022
2022
Later among the works it cites.
Y. Zhang, H. Jiang, Y. Miura, C. D. Manning, and C. P. Langlotz, “Contrastive learning of medical visual representations from paired images and text,” in Machine Learning for Healthcare Conference , 2022
2022
Later among the works it cites.
E. Tiu, E. Talius, P. Patel, C. P. Langlotz, A. Y. Ng, and P. Rajpurkar, “Expert-level detection of pathologies from unannotated chest x-ray images via self-supervised learning,” Nature Biomedical Engineering , 2022
2022
Later among the works it cites.
Z. Wang, Z. Wu, D. Agarwal, and J. Sun, “Medclip: Contrastive learning from unpaired medical images and text,” EMNLP , 2022
2022
Later among the works it cites.
A. Taleb, M. Kirchler, R. Monti, and C. Lippert, “Contig: Self-supervised multimodal contrastive learning for medical imaging with genetics,” in CVPR , 2022
2022
Later among the works it cites.
Y. Wang, C. M. Albrecht, N. A. A. Braham, L. Mou, and X. X. Zhu, “Self-supervised learning in remote sensing: A review,” Geoscience and Remote Sensing Magazine , 2022
2022
Later among the works it cites.
Y. Chen and L. Bruzzone, “Self-supervised sar-optical data fusion of sentinel-1/-2 images,” Transactions on Geoscience and Remote Sensing , 2022
2022
Later among the works it cites.
C. Schuhmann, R. Beaumont, R. Vencu, C. W. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman et al. , “Laion-5b: An open large-scale dataset for training next generation image-text models,” NeurIPS Datasets and Benchmarks Track , 2022
2022
Later among the works it cites.
Q. Cui, B. Zhou, Y. Guo, W. Yin, H. Wu, O. Yoshie, and Y. Chen, “Contrastive vision-language pre-training with limited resources,” in ECCV , 2022
2022
Later among the works it cites.
B. Sorscher, R. Geirhos, S. Shekhar, S. Ganguli, and A. S. Morcos, “Beyond neural scaling laws: beating power law scaling via data pruning,” NeurIPS , 2022
2022
Later among the works it cites.
Q. Zhang, S. Zuo, C. Liang, A. Bukharin, P. He, W. Chen, and T. Zhao, “Platon: Pruning large transformer models with upper confidence bound of weight importance,” in ICML , 2022
2022
Later among the works it cites.
Y. Li, F. Liang, L. Zhao, Y. Cui, W. Ouyang, J. Shao, F. Yu, and J. Yan, “Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm,” in ICLR , 2022
2022
Later among the works it cites.
H. Song, M. Kim, D. Park, Y. Shin, and J.-G. Lee, “Learning from noisy labels with deep neural networks: A survey,” TNNLS , 2022
2022
Later among the works it cites.
T. Srinivasan and Y. Bisk, “Worst of both worlds: Biases compound in pre-trained vision-and-language models,” in Workshop on Gender Bias in Natural Language Processing , 2022
2022
Later among the works it cites.
W. Liang, Y. Zhang, Y. Kwon, S. Yeung, and J. Zou, “Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning,” in NeurIPS , 2022
2022
Later among the works it cites.
A. Baevski, W.-N. Hsu, Q. Xu, A. Babu, J. Gu, and M. Auli, “data2vec: A general framework for self-supervised learning in speech, vision and language,” in ICML , 2022
2022
Later among the works it cites.
J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, E. H. Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus, “Emergent abilities of large language models,” TMLR , 2022
2022
Later among the works it cites.
Z. Chen, D. F. Fouhey, and A. Owens, “Sound localization by self-supervised time delay estimation,” in ECCV , 2022
2022
Later among the works it cites.
P. Xu, X. Zhu, and D. A. Clifton, “Multimodal learning with transformers: A survey,” T-PAMI , 2023
2023
Closest in time.
W. Wang, H. Bao, L. Dong, J. Bjorck, Z. Peng, Q. Liu, K. Aggarwal, O. Mohammed, S. Singhal, S. Som, and F. Wei, “Image as a foreign language: Beit pretraining for all vision and vision-language tasks,” CVPR , 2023
2023
Closest in time.
F. Zhan, Y. Yu, R. Wu, J. Zhang, and S. Lu, “Multimodal image synthesis and editing: A survey and taxonomy,” T-PAMI , 2023
2023
Closest in time.
L. Yang, Z. Zhang, S. Hong, R. Xu, Y. Zhao, Y. Shao, W. Zhang, M.-H. Yang, and B. Cui, “Diffusion models: A comprehensive survey of methods and applications,” ACM Computing Surveys , 2023
2023
Closest in time.
M. Tschannen, B. Mustafa, and N. Houlsby, “Image-and-language understanding from pixels only,” CVPR , 2023
2023
Closest in time.
Y. Li, H. Fan, R. Hu, C. Feichtenhofer, and K. He, “Scaling language-image pre-training via masking,” CVPR , 2023
2023
Closest in time.
F. Bordes, R. Balestriero, and P. Vincent, “Towards democratizing joint-embedding self-supervised learning,” arXiv , 2023
2023
Closest in time.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” NAACL , 2023
2023
Closest in time.
J. Lu, C. Clark, R. Zellers, R. Mottaghi, and A. Kembhavi, “Unified-io: A unified model for vision, language, and multi-modal tasks,” ICLR , 2023
2023
Closest in time.
Y. Gong, A. Rouditchenko, A. H. Liu, D. Harwath, L. Karlinsky, H. Kuehne, and J. R. Glass, “Contrastive audio-visual masked autoencoder,” in ICLR , 2023
2023
Closest in time.
P.-Y. Huang, V. Sharma, H. Xu, C. Ryali, Y. Li, S.-W. Li, G. Ghosh, J. Malik, and C. Feichtenhofer, “Mavil: Masked audio-video learners,” NuerIPS , 2023
2023
Closest in time.
J. Merullo, L. Castricato, C. Eickhoff, and E. Pavlick, “Linearly mapping from image to text space,” ICLR , 2023
2023
Closest in time.
J. Y. Koh, R. Salakhutdinov, and D. Fried, “Grounding language models to images for multimodal inputs and outputs,” ICML , 2023
2023
Closest in time.
J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” ICML , 2023
2023
Closest in time.
Y. Yu, J. Chung, H. Yun, J. Hessel, J. S. Park, X. Lu, R. Zellers, P. Ammanabrolu, R. Le Bras, G. Kim et al. , “Fusing pre-trained language models with multimodal prompts through reinforcement learning,” in CVPR , 2023
2023
Closest in time.
L. A. Passos, J. P. Papa, J. Del Ser, A. Hussain, and A. Adeel, “Multimodal audio-visual information fusion using canonical-correlated graph neural network for energy-efficient speech enhancement,” Information Fusion , 2023
2023
Closest in time.
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar et al. , “Llama: Open and efficient foundation language models,” arXiv , 2023
2023
Closest in time.
S. Yin, C. Fu, S. Zhao, K. Li, X. Sun, T. Xu, and E. Chen, “A survey on multimodal large language models,” arXiv , 2023
2023
Closest in time.
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu et al. , “Palm-e: An embodied multimodal language model,” in ICML , 2023
2023
Closest in time.
M. Yuksekgonul, F. Bianchi, P. Kalluri, D. Jurafsky, and J. Zou, “When and why vision-language models behave like bags-of-words, and what to do about it?” in ICLR , 2023
2023
Closest in time.
B. McKinzie, V. Shankar, J. Y. Cheng, Y. Yang, J. Shlens, and A. T. Toshev, “Robustness in multimodal learning under train-test modality mismatch,” ICML , 2023
2023
Closest in time.
N. Carlini, J. Hayes, M. Nasr, M. Jagielski, V. Sehwag, F. Tramer, B. Balle, D. Ippolito, and E. Wallace, “Extracting training data from diffusion models,” in USENIX Security Symposium , 2023
2023
Closest in time.
P. Rust, J. F. Lotz, E. Bugliarello, E. Salesky, M. de Lhoneux, and D. Elliott, “Language modelling with pixels,” in ICLR , 2023
2023
Closest in time.
A. Madani, B. Krause, E. R. Greene, S. Subramanian, B. P. Mohr, J. M. Holton, J. L. Olmos Jr, C. Xiong, Z. Z. Sun, R. Socher et al. , “Large language models generate functional protein sequences across diverse families,” Nature Biotechnology , 2023
2023
Closest in time.
R. Girdhar, A. El-Nouby, Z. Liu, M. Singh, K. V. Alwala, A. Joulin, and I. Misra, “Imagebind: One embedding space to bind them all,” in CVPR , 2023
2023
Closest in time.
R. Balestriero, M. Ibrahim, V. Sobal, A. Morcos, S. Shekhar, T. Goldstein, F. Bordes, A. Bardes, G. Mialon, Y. Tian et al. , “A cookbook of self-supervised learning,” arXiv , 2023
2023
Closest in time.
OpenAI, “Introducing chatgpt,” https://openai.com/blog/chatgpt , 2023
2023
Closest in time.
K. Lee, M. Joshi, I. Turc, H. Hu, F. Liu, J. Eisenschlos, U. Khandelwal, P. Shaw, M.-W. Chang, and K. Toutanova, “Pix2struct: Screenshot parsing as pretraining for visual language understanding,” ICML , 2023
2023
Closest in time.
P. P. Liang, A. Zadeh, and L.-P. Morency, “Foundations and recent trends in multimodal machine learning: Principles, challenges, and open questions,” ACM Computing Surveys , 2024
2024
Closest in time.
R. Shwartz Ziv and Y. LeCun, “To compress or not to compress—self-supervised learning and information theory: A review,” Entropy , 2024
2024
Closest in time.
Y. Ge, S. Zhao, Z. Zeng, Y. Ge, C. Li, X. Wang, and Y. Shan, “Making llama see and draw with seed tokenizer,” ICLR , 2024
2024
Closest in time.
Q. Sun, Q. Yu, Y. Cui, F. Zhang, X. Zhang, Y. Wang, H. Gao, J. Liu, T. Huang, and X. Wang, “Emu: Generative pretraining in multimodality,” in ICLR , 2024
2024
Closest in time.