Fetching the paper…
Reading the bibliography…
The emergence of Vision-Language Models (VLMs) is a significant advancement in integrating computer vision with Large Language Models (LLMs) to enhance multi-modal machine learning capabilities.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
On Evaluating Adversarial Robustness
Carlini, N.; Athalye, A.; Papernot, N.; Brendel, W.; Rauber, J.; Tsipras, D.; Goodfellow, I.; Madry, A.; and Kurakin, A. 2019 · 1902
Earlier work this paper cites.
Lewis, M.; Liu, Y.; Goyal, N.; Ghazvininejad, M.; Mohamed, A.; Levy, O.; Stoyanov, V.; and Zettlemoyer, L. 2019 · 1910
Earlier work this paper cites.
Estimation of the mean of a multivariate normal distribution
Stein, C. M. 1981 · 1981
Earlier work this paper cites.
Analysis of different norms and corresponding Lipschitz constants for global optimization
Paulavičius, R.; and Žilinskas, J. 2006 · 2006
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Gehman, S.; Gururangan, S.; Sap, M.; Choi, Y.; and Smith, N. A. 2020 · 2009
Earlier work this paper cites.
Denoising Diffusion Implicit Models
Song, J.; Meng, C.; and Ermon, S. 2022 · 2010
Earlier work this paper cites.
Evaluating the robustness of neural networks: An extreme value theory approach
Weng, T.-W.; Zhang, H.; Chen, P.-Y.; Yi, J.; Su, D.; Gao, Y.; Hsieh, C.-J.; and Daniel, L. 2018 · 2018
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Askell, A.; Bai, Y.; Chen, A.; Drain, D.; Ganguli, D.; Henighan, T.; Jones, A.; Joseph, N.; Mann, B.; DasSarma, N.; et al. 2021 · 2021
Earlier work this paper cites.
High-Resolution Image Synthesis with Latent Diffusion Models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2021 · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y.; Jones, A.; Ndousse, K.; Askell, A.; Chen, A.; DasSarma, N.; Drain, D.; Fort, S.; Ganguli, D.; Henighan, T.; et al. 2022 · 2022
Cited alongside, same era.
Li, J.; Li, D.; Xiong, C.; and Hoi, S. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C. L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P.; Leike, J.; and Lowe, R. 2022 · 2022
Cited alongside, same era.
DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models
Gong, S.; Li, M.; Feng, J.; Wu, Z.; and Kong, L. 2023 · 2023
Later among the works it cites.
Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023 · 2023
Later among the works it cites.
GREAT Score: Global Robustness Evaluation of Adversarial Perturbation using Generative Models
Li, Z.; Chen, P.-Y.; and Ho, T.-Y. 2023 · 2023
Later among the works it cites.
GPT-4v: Multimodal Language Model
OpenAI. 2023 · 2023
Later among the works it cites.
EVA-CLIP: Improved Training Techniques for CLIP at Scale
Sun, Q.; Fang, Y.; Wu, L.; Wang, X.; and Cao, Y. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Finetuned Language Models Are Zero-Shot Learners
Wei, J.; Bosma, M.; Zhao, V. Y.; Guu, K.; Yu, A. W.; Lester, B.; Du, N.; Dai, A. M.; and Le, Q. V. 2022 · 2022
Cited alongside, same era.
Are aligned neural networks adversarially aligned?
Carlini, N.; Nasr, M.; Choquette-Choo, C. A.; Jagielski, M.; Gao, I.; Awadalla, A.; Koh, P. W.; Ippolito, D.; Lee, K.; Tramer, F.; et al. 2023 · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; et al. 2023 · 2023
Cited alongside, same era.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Dai, W.; Li, J.; Li, D.; Tiong, A. M. H.; Zhao, J.; Wang, W.; Li, B.; Fung, P.; and Hoi, S. 2023 · 2023
Cited alongside, same era.
Eva: Exploring the limits of masked visual representation learning at scale
Fang, Y.; Wang, W.; Xie, B.; Sun, Q.; Wu, L.; Wang, X.; Huang, T.; Wang, X.; and Cao, Y. 2023 · 2023
Cited alongside, same era.
Visual Instruction Tuning
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023a
Cited in the paper.
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Liu, X.; Xu, N.; Chen, M.; and Xiao, C. 2023b
Cited in the paper.
Visual Adversarial Examples Jailbreak Aligned Large Language Models
Qi, X.; Huang, K.; Panda, A.; Henderson, P.; Wang, M.; and Mittal, P. 2023a
Cited in the paper.
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, G.; Anil, R.; Borgeaud, S.; Wu, Y.; Alayrac, J.-B.; Yu, J.; Soricut, R.; Schalkwyk, J.; Dai, A. M.; Hauth, A.; et al. 2023 · 2023
Later among the works it cites.
On Evaluating Adversarial Robustness of Large Vision-Language Models
Zhao, Y.; Pang, T.; Du, C.; Yang, X.; Li, C.; Cheung, N.-M.; and Lin, M. 2023 · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023 · 2023
Later among the works it cites.
Universal and Transferable Adversarial Attacks on Aligned Language Models
Zou, A.; Wang, Z.; Kolter, J. Z.; and Fredrikson, M. 2023 · 2023
Later among the works it cites.