Fetching the paper…
Reading the bibliography…
Vision-language models (VLMs) have improved significantly in their capabilities, but their complex architecture makes their safety alignment challenging.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. M. and Langford, J · 2002
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
Peters, J. and Schaal, S · 2008
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Goodfellow, I. J., Shlens, J., and Szegedy, C · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2016
Earlier work this paper cites.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P · 2017
Earlier work this paper cites.
Towards evaluating the robustness of neural networks
Carlini, N. and Wagner, D · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Adversarial examples are not bugs, they are features
Ilyas, A., Santurkar, S., Tsipras, D., Engstrom, L., Tran, B., and Madry, A · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Multi-exit vision transformer for dynamic inference
Bakhtiarnia, A., Zhang, Q., and Iosifidis, A · 2021
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Earlier work this paper cites.
Debertav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing, 2021
He, P., Gao, J., and Chen, W · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al · 2022
Cited alongside, same era.
A new generation of perspective api: Efficient multilingual character-level transformers
Lees, A., Tran, V. Q., Tay, Y., Sorensen, J., Gupta, J., Metzler, D., and Vasserman, L · 2022
Cited alongside, same era.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J., Li, D., Xiong, C., and Hoi, S · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Selfie: Self-interpretation of large language model embeddings
Chen, H., Vondrick, C., and Mao, C · 2024
Closest in time.
Dola: Decoding by contrasting layers improves factuality in large language models
Chuang, Y.-S., Xie, Y., Luo, H., Kim, Y., Glass, J. R., and He, P · 2024
Closest in time.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Closest in time.
Onellm: One framework to align all modalities with language
Han, J., Gong, K., Zhang, Y., Wang, J., Zhang, K., Lin, D., Qiao, Y., Gao, P., and Yue, X · 2024
Closest in time.
Intrinsic evaluation of unlearning using parametric knowledge traces
Hong, Y., Yu, L., Yang, H., Ravfogel, S., and Geva, M · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Figstep: Jailbreaking large vision-language models via typographic visual prompts
Gong, Y., Ran, D., Liu, J., Wang, C., Cong, T., Wang, A., Duan, S., and Wang, X · 2023
Cited alongside, same era.
Llama guard: Llm-based input-output safeguard for human-ai conversations, 2023
Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., and Khabsa, M · 2023
Cited alongside, same era.
Beavertails: Towards improved safety alignment of LLM via a human-preference dataset
Ji, J., Liu, M., Dai, J., Pan, X., Zhang, C., Bian, C., Chen, B., Sun, R., Wang, Y., and Yang, Y · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2023
Cited alongside, same era.
On the adversarial robustness of multi-modal foundation models
Schlarmann, C. and Hein, M · 2023
Cited alongside, same era.
Dynamic inference with grounding based vision and language models
Uzkent, B., Garg, A., Zhu, W., Doshi, K., Yi, J., Wang, X., and Omar, M · 2023
Cited alongside, same era.
Lgvit: Dynamic early exiting for accelerating vision transformer
Xu, G., Hao, J., Shen, L., Hu, H., Luo, Y., Lin, H., and Shen, J · 2023
Cited alongside, same era.
Closest in time.
VLFeedback: A large-scale AI feedback dataset for large vision-language models alignment
Li, L., Xie, Z., Li, M., Chen, S., Wang, P., Chen, L., Yang, Y., Wang, B., Kong, L., and Liu, Q · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024a
Liu, H., Li, C., Li, Y., Li, B., Zhang, Y., Shen, S., and Lee, Y. J · 2024
Closest in time.
Jailbreakv: A benchmark for assessing the robustness of multimodal large language models against jailbreak attacks
Luo, W., Ma, S., Liu, X., Guo, X., and Xiao, C · 2024
Closest in time.
Jailbreaking attack against multimodal large language model
Niu, Z., Ren, H., Gao, X., Hua, G., and Jin, R · 2024
Closest in time.
Strengthening multimodal large language model with bootstrapped preference optimization
Pi, R., Han, T., Xiong, W., Zhang, J., Liu, R., Pan, R., and Zhang, T · 2024
Closest in time.
XSTest: A test suite for identifying exaggerated safety behaviours in large language models
Röttger, P., Kirk, H., Vidgen, B., Attanasio, G., Bianchi, F., and Hovy, D · 2024
Closest in time.
Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models
Shayegani, E., Dong, Y., and Abu-Ghazaleh, N · 2024
Closest in time.
Aligning large multimodal models with factually augmented RLHF
Sun, Z., Shen, S., Cao, S., Liu, H., Li, C., Shen, Y., Gan, C., Gui, L., Wang, Y.-X., Yang, Y., Keutzer, K., and Darrell, T · 2024
Closest in time.
Jailbroken: How does llm safety training fail?
Wei, A., Haghtalab, N., and Steinhardt, J · 2024
Closest in time.
Rewards-in-context: Multi-objective alignment of foundation models with dynamic preference adjustment
Yang, R., Pan, X., Luo, F., Qiu, S., Zhong, H., Yu, D., and Chen, J · 2024
Closest in time.
Spa-vl: A comprehensive safety preference alignment dataset for vision language model, 2024
Zhang, Y., Chen, L., Zheng, G., Gao, Y., Zheng, R., Fu, J., Yin, Z., Jin, S., Qiao, Y., Huang, X., Zhao, F., Gui, T., and Shao, J · 2024
Closest in time.
Accelerating greedy coordinate gradient and general prompt optimization via probe sampling
Zhao, Y., Zheng, W., Cai, T., Long, D. X., Kawaguchi, K., Goyal, A., and Shieh, M · 2024
Closest in time.
Aligning modalities in vision large language models via preference fine-tuning, 2024
Zhou, Y., Cui, C., Rafailov, R., Finn, C., and Yao, H · 2024
Closest in time.
Layer by layer: Uncovering hidden representations in language models, 2025
Skean, O., Arefin, M. R., Zhao, D., Patel, N., Naghiyev, J., LeCun, Y., and Shwartz-Ziv, R · 2025
Closest in time.