Fetching the paper…
Reading the bibliography…
The rise of multimodal large language models has introduced innovative human-machine interaction paradigms but also significant challenges in machine learning safety.
Universal adversarial triggers for attacking and analyzing nlp, 2021
Wallace, E., Feng, S., Kandpal, N., Gardner, M., and Singh, S · 1908
Earlier work this paper cites.
Language models are few-shot learners, 2020
Brown, T. B · 2005
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models, 2020
Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A · 2009
Earlier work this paper cites.
Evasion Attacks against Machine Learning at Test Time , pp. 387–402
Biggio, B., Corona, I., Maiorca, D., Nelson, B., Šrndić, N., Laskov, P., Giacinto, G., and Roli, F · 2013
Earlier work this paper cites.
Intriguing properties of neural networks, 2014
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R · 2014
Earlier work this paper cites.
Adversarial examples for evaluating reading comprehension systems
Jia, R. and Liang, P · 2017
Earlier work this paper cites.
Multimodal machine learning: A survey and taxonomy
Baltrušaitis, T., Ahuja, C., and Morency, L.-P · 2018
Earlier work this paper cites.
Audio adversarial examples: Targeted attacks on speech-to-text, 2018
Carlini, N. and Wagner, D · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Hotflip: White-box adversarial examples for text classification, 2018
Ebrahimi, J., Rao, A., Lowd, D., and Dou, D · 2018
Earlier work this paper cites.
Physical adversarial examples for object detectors, 2018
Eykholt, K., Evtimov, I., Fernandes, E., Li, B., Rahmati, A., Tramer, F., Prakash, A., Kohno, T., and Song, D · 2018
Earlier work this paper cites.
Fooling end-to-end speaker verification with adversarial examples
Kreuk, F., Adi, Y., Cissé, M., and Keshet, J · 2018
Earlier work this paper cites.
Adversarial attacks against automatic speech recognition systems via psychoacoustic hiding, 2018
Schönherr, L., Kohls, K., Zeiler, S., Holz, T., and Kolossa, D · 2018
Earlier work this paper cites.
Detoxify
Hanu, L. and Unitary team · 2020
Earlier work this paper cites.
Universal adversarial attacks on spoken language assessment systems
Raina, V., Gales, M., and Knill, K · 2020
Earlier work this paper cites.
Ethical and social risks of harm from language models, 2021
Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P.-S., Cheng, M., Glaese, M., Balle, B., Kasirzadeh, A., Kenton, Z., Brown, S., Hawkins, W., Stepleton, T., Biles, C., Birhane, A., Haas, J., Rimell, L., Hendricks, L. A., Isaac, W., Legassick, S., Irving, G., and Gabriel, I · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning, 2022
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., Ring, R., Rutherford, E., Cabi, S., Han, T., Gong, Z., Samangooei, S., Monteiro, M., Menick, J., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Binkowski, M., Barreira, R., Vinyals, O., Zisserman, A., and Simonyan, K · 2022
Earlier work this paper cites.
On the opportunities and risks of foundation models, 2022
Bommasani, R · 2022
Earlier work this paper cites.
Beats: Audio pre-training with acoustic tokenizers, 2022
Chen, S., Wu, Y., Wang, C., Liu, S., Tompkins, D., Chen, Z., and Wei, F · 2022
Earlier work this paper cites.
Physical adversarial attack on a robotic arm
Jia, Y., Poskitt, C. M., Sun, J., and Chattopadhyay, S · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback, 2022
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Cited alongside, same era.
Red teaming language models with language models, 2022
Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G · 2022
Cited alongside, same era.
Robust speech recognition via large-scale weak supervision, 2022
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I · 2022
Cited alongside, same era.
Artificial Intelligence and the Problem of Control , pp. 19–24
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., de las Casas, D., Hanna, E. B., Bressand, F., Lengyel, G., Bour, G., Lample, G., Lavaud, L. R., Saulnier, L., Lachaux, M.-A., Stock, P., Subramanian, S., Yang, S., Antoniak, S., Scao, T. L., Gervet, T., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2024
Later among the works it cites.
Advwave: Stealthy adversarial jailbreak attack against large audio-language models, 2024
Kang, M., Xu, C., and Li, B · 2024
Later among the works it cites.
Towards efficient visual-language alignment of the q-former for visual reasoning tasks, 2024
Kim, S., Lee, A., Park, J., Chung, A., Oh, J., and Lee, J.-Y · 2024
Later among the works it cites.
Kumar, V., Liao, Z., Jones, J., and Sun, H · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Russell, S · 2022
Cited alongside, same era.
Finetuned language models are zero-shot learners, 2022
Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2022
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Cited alongside, same era.
Deep reinforcement learning from human preferences, 2023
Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D · 2023
Cited alongside, same era.
Chu, Y., Xu, J., Zhou, X., Yang, Q., Zhang, S., Yan, Z., Zhou, C., and Zhou, J · 2023
Cited alongside, same era.
Advddos: Zero-query adversarial attacks against commercial speech recognition systems
Ge, Y., Zhao, L., Wang, Q., Duan, Y., and Du, M · 2023
Cited alongside, same era.
Voice biometrics fusion for enhanced security and speaker recognition: A comprehensive review
Koffi, E · 2023
Cited alongside, same era.
Visual adversarial examples jailbreak aligned large language models, 2023
Qi, X., Huang, K., Panda, A., Henderson, P., Wang, M., and Mittal, P · 2023
Cited alongside, same era.
Li, Y., Guo, H., Zhou, K., Zhao, W. X., and Wen, J.-R · 2024
Later among the works it cites.
Autodan: Generating stealthy jailbreak prompts on aligned large language models, 2024
Liu, X., Xu, N., Chen, M., and Xiao, C · 2024
Later among the works it cites.
Jailbreaking prompt attack: A controllable adversarial attack against diffusion models, 2024
Ma, J., Cao, A., Xiao, Z., Li, Y., Zhang, J., Ye, C., and Zhao, J · 2024
Later among the works it cites.
User interaction patterns and breakdowns in conversing with llm-powered voice assistants
Mahmood, A., Wang, J., Yao, B., Wang, D., and Huang, C.-M · 2024
Later among the works it cites.
Rule based rewards for language model safety, 2024
Mu, T., Helyar, A., Heidecke, J., Achiam, J., Vallone, A., Kivlichan, I., Lin, M., Beutel, A., Schulman, J., and Weng, L · 2024
Later among the works it cites.
OpenAI · 2024
Later among the works it cites.
Muting whisper: A universal acoustic adversarial attack on speech foundation models
Raina, V., Ma, R., McGhee, C., Knill, K., and Gales, M · 2024
Later among the works it cites.
Failures to find transferable image jailbreaks between vision-language models, 2024
Schaeffer, R., Valentine, D., Bailey, L., Chua, J., Eyzaguirre, C., Durante, Z., Benton, J., Miranda, B., Sleight, H., Hughes, J., Agrawal, R., Sharma, M., Emmons, S., Koyejo, S., and Perez, E · 2024
Later among the works it cites.
Salmonn: Towards generic hearing abilities for large language models, 2024
Tang, C., Yu, W., Sun, G., Chen, X., Tan, T., Li, W., Lu, L., Ma, Z., and Zhang, C · 2024
Later among the works it cites.
A comprehensive study of jailbreak attack versus defense for large language models, 2024
Xu, Z., Liu, Y., Deng, G., Li, Y., and Picek, S · 2024
Later among the works it cites.
Audio is the achilles’ heel: Red teaming audio large multimodal models, 2024
Yang, H., Qu, L., Shareghi, E., and Haffari, G · 2024
Later among the works it cites.
Jailbreak attacks and defenses against large language models: A survey, 2024
Yi, S., Liu, Y., Sun, Z., Cong, T., He, X., Song, J., Xu, K., and Li, Q · 2024
Later among the works it cites.
Jailbreak vision language models via bi-modal adversarial prompt, 2024
Ying, Z., Liu, A., Zhang, T., Yu, Z., Liang, S., Liu, X., and Tao, D · 2024
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models, 2023
Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., and Fredrikson, M · 2024
Later among the works it cites.