Fetching the paper…
Reading the bibliography…
Multimodal learning is defined as learning over multiple heterogeneous input modalities such as video, audio, and text.
Benchmarking neural network robustness to common corruptions and perturbations, 2019
Hendrycks, D. and Dietterich, T · 1903
Earlier work this paper cites.
Using videos to evaluate image model robustness, 2019
Gu, K., Yang, B., Ngiam, J., Le, Q., and Shlens, J · 1904
Earlier work this paper cites.
A training algorithm for optimal margin classifiers
Boser, B. E., Guyon, I. M., and Vapnik, V. N · 1992
Earlier work this paper cites.
Self-supervised multimodal versatile networks
Alayrac, J., Recasens, A., Schneider, R., Arandjelovic, R., Ramapuram, J., Fauw, J. D., Smaira, L., Dieleman, S., and Zisserman, A · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M. and Hyvärinen, A · 2010
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., Dean, J., et al · 2015
Earlier work this paper cites.
Deep learning , volume 1
Goodfellow, I., Bengio, Y., Courville, A., and Bengio, Y · 2016
Earlier work this paper cites.
Deep networks with stochastic depth, 2016
Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K · 2016
Earlier work this paper cites.
Improved deep metric learning with multi-class n-pair loss objective
Sohn, K · 2016
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F., Ellis, D. P., Freedman, D., Jansen, A., Lawrence, W., Moore, R. C., Plakal, M., and Ritter, M · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., et al · 2017
Earlier work this paper cites.
Learning without forgetting
Li, Z. and Hoiem, D · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A · 2017
Earlier work this paper cites.
icarl: Incremental classifier and representation learning
Rebuffi, S.-A., Kolesnikov, A., Sperl, G., and Lampert, C. H · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Generating natural language adversarial examples
Alzantot, M., Sharma, Y., Elgohary, A., Ho, B.-J., Srivastava, M., and Chang, K.-W · 2018
Earlier work this paper cites.
Audio adversarial examples: Targeted attacks on speech-to-text
Carlini, N. and Wagner, D · 2018
Earlier work this paper cites.
End-to-end incremental learning
Castro, F. M., Marín-Jiménez, M. J., Guil, N., Schmid, C., and Alahari, K · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Generalisation in humans and deep neural networks
Geirhos, R., Temme, C. R. M., Rauber, J., Schütt, H. H., Bethge, M., and Wichmann, F. A · 2018
Cited alongside, same era.
Lifelong learning via progressive distillation and retrospection
Hou, S., Pan, X., Loy, C. C., Wang, Z., and Lin, D · 2018
Cited alongside, same era.
Learn to combine modalities in multimodal deep learning, 2018
Liu, K., Li, Y., Xu, N., and Natarajan, P · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Cited alongside, same era.
Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models
Barbu, A., Mayo, D., Alverio, J., Luo, W., Wang, C., Gutfreund, D., Tenenbaum, J., and Katz, B · 2019
Cited alongside, same era.
Florence: A new foundation model for computer vision
Yuan, L., Chen, D., Chen, Y.-L., Codella, N., Dai, X., Gao, J., Hu, H., Huang, X., Li, B., Li, C., et al · 2021
Later among the works it cites.
mixup: Beyond empirical risk minimization, 2017
Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D · 2021
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Later among the works it cites.
Mae-ast: Masked autoencoding audio spectrogram transformer
Baade, A., Peng, P., and Harwath, D · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Unlabeled data improves adversarial robustness
Carmon, Y., Raghunathan, A., Schmidt, L., Duchi, J. C., and Liang, P. S · 2019
Cited alongside, same era.
SpecAugment: A simple data augmentation method for automatic speech recognition
Park, D. S., Chan, W., Zhang, Y., Chiu, C.-C., Zoph, B., Cubuk, E. D., and Le, Q. V · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Cited alongside, same era.
Do imagenet classifiers generalize to imagenet?
Recht, B., Roelofs, R., Schmidt, L., and Shankar, V · 2019
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Cited alongside, same era.
Does my multimodal model learn cross-modal interactions? it’s harder to tell than you might think!
Hessel, J. and Lee, L · 2020
Cited alongside, same era.
Measuring robustness to natural distribution shifts in image classification
Taori, R., Dave, A., Shankar, V., Carlini, N., Recht, B., and Schmidt, L · 2020
Cited alongside, same era.
Carlini, N., Tramer, F., Dvijotham, K., and Kolter, J. Z · 2022
Later among the works it cites.
Masked spectrogram prediction for self-supervised audio pre-training, 2022
Chong, D., Wang, H., Zhou, P., and Zeng, Q · 2022
Later among the works it cites.
Masked autoencoders as spatiotemporal learners
Feichtenhofer, C., Fan, H., Li, Y., and He, K · 2022
Later among the works it cites.
Multimodal masked autoencoders learn transferable representations
Geng, X., Liu, H., Lee, L., Schuurams, D., Levine, S., and Abbeel, P · 2022
Later among the works it cites.
Audiovisual masked autoencoders
Georgescu, M.-I., Fonseca, E., Ionescu, R. T., Lucic, M., Schmid, C., and Arnab, A · 2022
Later among the works it cites.
Omnimae: Single model masked pretraining on images and videos
Girdhar, R., El-Nouby, A., Singh, M., Alwala, K. V., Joulin, A., and Misra, I · 2022
Later among the works it cites.
Contrastive audio-visual masked autoencoder
Gong, Y., Rouditchenko, A., Liu, A. H., Harwath, D., Karlinsky, L., Kuehne, H., and Glass, J · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2022
Later among the works it cites.
Masked autoencoders that listen, 2022
Huang, P.-Y., Xu, H., Li, J., Baevski, A., Auli, M., Galuba, W., Metze, F., and Feichtenhofer, C · 2022
Later among the works it cites.
Patching open-vocabulary models by interpolating weights, 2022
Ilharco, G., Wortsman, M., Gadre, S. Y., Song, S., Hajishirzi, H., Kornblith, S., Farhadi, A., and Schmidt, L · 2022
Later among the works it cites.
Highmmt: Towards modality and task generalization for high-modality representation learning
Liang, P. P., Lyu, Y., Fan, X., Mo, S., Yogatama, D., Morency, L.-P., and Salakhutdinov, R · 2022
Later among the works it cites.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Tong, Z., Song, Y., Wang, J., and Wang, L · 2022
Later among the works it cites.
Image as a foreign language: Beit pretraining for all vision and vision-language tasks
Wang, W., Bao, H., Dong, L., Bjorck, J., Peng, Z., Liu, Q., Aggarwal, K., Mohammed, O. K., Singhal, S., Som, S., et al · 2022
Later among the works it cites.
Robust fine-tuning of zero-shot models
Wortsman, M., Ilharco, G., Kim, J. W., Li, M., Kornblith, S., Roelofs, R., Lopes, R. G., Hajishirzi, H., Farhadi, A., Namkoong, H., and Schmidt, L · 2022
Later among the works it cites.
Masked autoencoders that listen
Xu, H., Li, J., Baevski, A., Auli, M., Galuba, W., Metze, F., Feichtenhofer, C., et al · 2022
Later among the works it cites.
i-code: An integrative and composable multimodal learning framework
Yang, Z., Fang, Y., Zhu, C., Pryzant, R., Chen, D., Shi, Y., Xu, Y., Qian, Y., Gao, M., Chen, Y.-L., et al · 2022
Later among the works it cites.
Learning visual representation from modality-shared contrastive language-image pre-training, 2022
You, H., Zhou, L., Xiao, B., Codella, N., Cheng, Y., Xu, R., Chang, S.-F., and Yuan, L · 2022
Later among the works it cites.