Fetching the paper…
Reading the bibliography…
Multi-modal foundation models like OpenFlamingo, LLaVA, and GPT-4 are increasingly used for various real-world tasks.
Caltech-256 object category dataset
Griffin, G., Holub, A., and Perona, P · 2007
Earlier work this paper cites.
Automated flower classification over a large number of classes
Nilsback, M.-E. and Zisserman, A · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning
Coates, A., Ng, A., and Lee, H · 2011
Earlier work this paper cites.
Cats and dogs
Parkhi, O. M., Vedaldi, A., Zisserman, A., and Jawahar, C. V · 2012
Earlier work this paper cites.
3d object representations for fine-grained categorization
Krause, J., Stark, M., Deng, J., and Fei-Fei, L · 2013
Earlier work this paper cites.
Fine-grained visual classification of aircraft, 2013
Maji, S., Rahtu, E., Kannala, J., Blaschko, M., and Vedaldi, A · 2013
Earlier work this paper cites.
Describing textures in the wild
Cimpoi, M., Maji, S., Kokkinos, I., Mohamed, S., and Vedaldi, A · 2014
Earlier work this paper cites.
Microsoft COCO: common objects in context
Lin, T., Maire, M., Belongie, S. J., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I. J., and Fergus, R · 2014
Earlier work this paper cites.
VQA: visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D · 2015
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Goodfellow, I. J., Shlens, J., and Szegedy, C · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J., and Lazebnik, S · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Vedantam, R., Zitnick, C. L., and Parikh, D · 2015
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
Adversarial examples for evaluating reading comprehension systems
Jia, R. and Liang, P · 2017
Earlier work this paper cites.
Hotflip: White-box adversarial examples for text classification
Ebrahimi, J., Rao, A., Lowd, D., and Dou, D · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2018
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A · 2018
Earlier work this paper cites.
Rotation equivariant cnns for digital pathology
Veeling, B. S., Linmans, J., Winkens, J., Cohen, T., and Welling, M · 2018
Earlier work this paper cites.
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
Helber, P., Bischke, B., Dengel, A., and Borth, D · 2019
Earlier work this paper cites.
Metric learning for adversarial robustness
Mao, C., Zhong, Z., Yang, J., Vondrick, C., and Ray, B · 2019
Cited alongside, same era.
Towards vqa models that can read
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M · 2019
Cited alongside, same era.
Learning robust global representations by penalizing local predictive power
Wang, H., Ge, S., Lipton, Z., and Xing, E. P · 2019
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. E · 2020
Cited alongside, same era.
Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks
Croce, F. and Hein, M · 2020
Cited alongside, same era.
Self-supervised adversarial robustness for the low-label, high-data regime
Gowal, S., Huang, P.-S., van den Oord, A., Mann, T., and Kohli, P · 2020
Shikra: Unleashing multimodal LLM’s referential dialogue magic
Chen, K., Zhang, Z., Zeng, W., Zhang, R., Zhu, F., and Zhao, R · 2023
Later among the works it cites.
Reproducible scaling laws for contrastive language-image learning
Cherti, M., Beaumont, R., Wightman, R., Wortsman, M., Ilharco, G., Gordon, C., Schuhmann, C., Schmidt, L., and Jitsev, J · 2023
Later among the works it cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Later among the works it cites.
How robust is google’s bard to adversarial image attacks?
Dong, Y., Chen, H., Chen, J., Fang, Z., Yang, X., Zhang, Y., Tian, Y., Su, H., and Zhu, J · 2023
Later among the works it cites.
Grounding language models to images for multimodal inputs and outputs
Koh, J. Y., Salakhutdinov, R., and Fried, D · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bootstrap your own latent-a new approach to self-supervised learning
Grill, J.-B., Strub, F., Altché, F., Tallec, C., Richemond, P., Buchatskaya, E., Doersch, C., Avila Pires, B., Guo, Z., Gheshlaghi Azar, M., et al · 2020
Cited alongside, same era.
Robust pre-training by adversarial contrastive learning
Jiang, Z., Chen, T., Chen, T., and Wang, Z · 2020
Cited alongside, same era.
Adversarial self-supervised contrastive learning
Kim, M., Tack, J., and Hwang, S. J · 2020
Cited alongside, same era.
When does contrastive learning preserve adversarial robustness from pretraining to finetuning?
Fan, L., Liu, S., Chen, P.-Y., Zhang, G., and Gan, C · 2021
Cited alongside, same era.
The many faces of robustness: A critical analysis of out-of-distribution generalization
Hendrycks, D., Basart, S., Mu, N., Kadavath, S., Wang, F., Dorundo, E., Desai, R., Zhu, T., Parajuli, S., Guo, M., et al · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Cited alongside, same era.
Later among the works it cites.
OBELICS: An open web-scale filtered dataset of interleaved image-text documents
Laurençon, H., Saulnier, L., Tronchon, L., Bekman, S., Singh, A., Lozhkov, A., Wang, T., Karamcheti, S., Rush, A. M., Kiela, D., Cord, M., and Sanh, V · 2023
Later among the works it cites.
Rethinking the effect of data augmentation in adversarial contrastive learning
Luo, R., Wang, Y., and Wang, Y · 2023
Later among the works it cites.
Understanding zero-shot adversarial robustness for large-scale models
Mao, C., Geng, S., Yang, J., Wang, X. E., and Vondrick, C · 2023
Later among the works it cites.
Introducing mpt-7b: A new standard for open-source, commercially usable LLMs, 2023
MosaicML · 2023
Later among the works it cites.
Visual adversarial examples jailbreak large language models
Qi, X., Huang, K., Panda, A., Wang, M., and Mittal, P · 2023
Later among the works it cites.
On the adversarial robustness of multi-modal foundation models
Schlarmann, C. and Hein, M · 2023
Later among the works it cites.
Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models
Shayegani, E., Dong, Y., and Abu-Ghazaleh, N · 2023
Later among the works it cites.
Shen, X., Chen, Z., Backes, M., Shen, Y., and Zhang, Y · 2023
Later among the works it cites.
Revisiting adversarial training for imagenet: Architectures, training and generalization across threat models
Singh, N. D., Croce, F., and Hein, M · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Later among the works it cites.
Enhancing adversarial contrastive learning via adversarial invariant regularization
Xu, X., Zhang, J., Liu, F., Sugiyama, M., and Kankanhalli, M. S · 2023
Later among the works it cites.
On evaluating adversarial robustness of large vision-language models
Zhao, Y., Pang, T., Du, C., Yang, X., Li, C., Cheung, N.-M., and Lin, M · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Later among the works it cites.
Agent smith: A single image can jailbreak one million multimodal llm agents exponentially fast
Gu, X., Zheng, X., Pang, T., Du, C., Liu, Q., Wang, Y., Jiang, J., and Lin, M · 2024
Closest in time.
Adversarial attack and defense in deep ranking
Zhou, M., Wang, L., Niu, Z., Zhang, Q., Zheng, N., and Hua, G · 2024
Closest in time.