Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have surged in popularity in recent months, but they have demonstrated concerning capabilities to generate harmful content when manipulated.
On evaluating adversarial robustness
Carlini, N., Athalye, A., Papernot, N., Brendel, W., Rauber, J., Tsipras, D., Goodfellow, I., Madry, A., and Kurakin, A · 1902
Earlier work this paper cites.
Subspace attack: Exploiting promising subspaces for query-efficient black-box attacks
Yan, Z., Guo, Y., and Zhang, C · 1906
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer, September 2023
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 1910
Earlier work this paper cites.
Convergence of a block coordinate descent method for nondifferentiable minimization
Tseng, P · 2001
Earlier work this paper cites.
Language models are few-shot learners, July 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2005
Earlier work this paper cites.
Towards evaluating the robustness of neural networks
Carlini, N. and Wagner, D · 2017
Earlier work this paper cites.
Boosting adversarial attacks with momentum
Dong, Y., Liao, F., Pang, T., Su, H., Zhu, J., Hu, X., and Li, J · 2018
Earlier work this paper cites.
Adversarial risk and the dangers of evaluating against weak attacks
Uesato, J., O’Donoghue, B., Kohli, P., and van den Oord, A · 2018
Earlier work this paper cites.
With great training comes great vulnerability: Practical attacks against transfer learning
Wang, B., Yao, Y., Viswanath, B., Zheng, H., and Zhao, B. Y · 2018
Earlier work this paper cites.
Improving black-box adversarial attacks with a transfer-based prior
Cheng, S., Dong, Y., Pang, T., Su, H., and Zhu, J · 2019
Earlier work this paper cites.
Stateful detection of black-box adversarial attacks
Chen, S., Carlini, N., and Wagner, D · 2020
Earlier work this paper cites.
The pile: An 800GB dataset of diverse text for language modeling, December 2020
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C · 2020
Earlier work this paper cites.
Black-box adversarial attack with transferable model-based embedding
Huang, Z. and Zhang, T · 2020
Earlier work this paper cites.
AutoPrompt: Eliciting knowledge from language models with automatically generated prompts
Shin, T., Razeghi, Y., Logan IV, R. L., Wallace, E., and Singh, S · 2020
Earlier work this paper cites.
Towards understanding and improving the transferability of adversarial examples in deep neural networks
Wu, L. and Zhu, Z · 2020
Earlier work this paper cites.
Model extraction and adversarial transferability, your BERT is vulnerable!
He, X., Lyu, L., Sun, L., and Xu, Q · 2021
Earlier work this paper cites.
Ethical and social risks of harm from language models
Weidinger, L., Mellor, J. F. J., Rauh, M., Griffin, C., Uesato, J., Huang, P.-S., Cheng, M., Glaese, M., Balle, B., Kasirzadeh, A., Kenton, Z., Brown, S. M., Hawkins, W. T., Stepleton, T., Biles, C., Birhane, A., Haas, J., Rimell, L., Hendricks, L. A., Isaac, W. S., Legassick, S., Irving, G., and Gabriel, I · 2021
Earlier work this paper cites.
Constitutional AI: Harmlessness from AI feedback, December 2022
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Chen, C., Olsson, C., Olah, C., Hernandez, D., Drain, D., Ganguli, D., Li, D., Tran-Johnson, E., Perez, E., Kerr, J., Mueller, J., Ladish, J., Landau, J., Ndousse, K., Lukosuite, K., Lovitt, L., Sellitto, M., Elhage, N., Schiefer, N., Mercado, N., DasSarma, N., Lasenby, R., Larson, R., Ringer, S., Johnston, S., Kravec, S., Showk, S. E., Fort, S., Lanham, T., Telleen-Lawton, T., Conerly, T., Henighan, T., Hume, T., Bowman, S. R., Hatfield-Dodds, Z., Mann, B., Amodei, D., Joseph, N., McCandlish, S., Brown, T., and Kaplan, J · 2022
Earlier work this paper cites.
Evaluating the susceptibility of pre-trained language models via handcrafted adversarial examples
Branch, H. J., Cefalu, J. R., McHugh, J., Hujer, L., Bahl, A., del Castillo Iglesias, D., Heichman, R., and Darwishi, R · 2022
Earlier work this paper cites.
Blackbox attacks via surrogate ensemble search
Cai, Z., Song, C., Krishnamurthy, S., Roy-Chowdhury, A., and Asif, S · 2022
Earlier work this paper cites.
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., Jones, A., Bowman, S., Chen, A., Conerly, T., DasSarma, N., Drain, D., Elhage, N., El-Showk, S., Fort, S., Hatfield-Dodds, Z., Henighan, T., Hernandez, D., Hume, T., Jacobson, J., Johnston, S., Kravec, S., Olsson, C., Ringer, S., Tran-Johnson, E., Amodei, D., Brown, T., Joseph, N., McCandlish, S., Olah, C., Kaplan, J., and Clark, J · 2022
Earlier work this paper cites.
Improving alignment of dialogue agents via targeted human judgements
Glaese, A., McAleese, N., Trębacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., et al · 2022
Cited alongside, same era.
Attacking deep networks with surrogate-based adversarial black-box methods is easy
Lord, N. A., Mueller, R., and Bertinetto, L · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback, March 2022
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Cited alongside, same era.
Red teaming language models with language models
Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G · 2022
Cited alongside, same era.
Ignore previous prompt: Attack techniques for language models
Pretraining language models with human preferences
Korbak, T., Shi, K., Chen, A., Bhalerao, R. V., Buckley, C., Phang, J., Bowman, S. R., and Perez, E · 2023
Later among the works it cites.
Open sesame! Universal black box jailbreaking of large language models, September 2023
Lapid, R., Langberg, R., and Sipper, M · 2023
Later among the works it cites.
ADDA: An adversarial direction-guided decision-based attack via multiple surrogate models
Li, W. and Liu, X · 2023
Later among the works it cites.
Autodan: Generating stealthy jailbreak prompts on aligned large language models, 2023
Liu, X., Xu, N., Chen, M., and Xiao, C · 2023
Later among the works it cites.
Improving adversarial transferability via model alignment, November 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Perez, F. and Ribeiro, I · 2022
Cited alongside, same era.
Optimization for Data Analysis
Recht, B. and Wright, S. J · 2022
Cited alongside, same era.
Taxonomy of risks posed by language models
Weidinger, L., Uesato, J., Rauh, M., Griffin, C., Huang, P.-S., Mellor, J. F. J., Glaese, A., Cheng, M., Balle, B., Kasirzadeh, A., Biles, C., Brown, S. M., Kenton, Z., Hawkins, W. T., Stepleton, T., Birhane, A., Hendricks, L. A., Rimell, L., Isaac, W. S., Haas, J., Legassick, S., Irving, G., and Gabriel, I · 2022
Cited alongside, same era.
Adversarial attacks on GPT-4 via simple random search, December 2023
Andriushchenko, M · 2023
Cited alongside, same era.
Are aligned neural networks adversarially aligned?, June 2023
Carlini, N., Nasr, M., Choquette-Choo, C. A., Jagielski, M., Gao, I., Awadalla, A., Koh, P. W., Ippolito, D., Lee, K., Tramer, F., and Schmidt, L · 2023
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries, October 2023
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E · 2023
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2023
Cited alongside, same era.
RedPajama: An open dataset for training large language models, October 2023
Computer, T · 2023
Cited alongside, same era.
Ma, A., Farahmand, A.-m., Pan, Y., Torr, P., and Gu, J · 2023
Later among the works it cites.
Tree of attacks: Jailbreaking black-box LLMs automatically, December 2023
Mehrotra, A., Zampetakis, M., Kassianik, P., Nelson, B., Anderson, H., Singer, Y., and Karbasi, A · 2023
Later among the works it cites.
Language model inversion, November 2023
Morris, J. X., Zhao, W., Chiu, J. T., Shmatikov, V., and Rush, A. M · 2023
Later among the works it cites.
Gpt-4 system card
OpenAI · 2023
Later among the works it cites.
Penedo, G., Malartic, Q., Hesslow, D., Cojocaru, R., Cappelli, A., Alobeidli, H., Pannier, B., Almazrouei, E., and Launay, J · 2023
Later among the works it cites.
Shah, M. A., Sharma, R., Dhamyal, H., Olivier, R., Shah, A., Konan, J., Alharthi, D., Bukhari, H. T., Baali, M., Deshmukh, S., Kuhlmann, M., Raj, B., and Singh, R · 2023
Later among the works it cites.
Shen, X., Chen, Z., Backes, M., Shen, Y., and Zhang, Y · 2023
Later among the works it cites.
Jailbroken: How does LLM safety training fail?, July 2023
Wei, A., Haghtalab, N., and Steinhardt, J · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K. R · 2023
Later among the works it cites.
Low-resource languages jailbreak gpt-4
Yong, Z.-X., Menghini, C., and Bach, S. H · 2023
Later among the works it cites.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts, 2023
Yu, J., Lin, X., Yu, Z., and Xing, X · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models, 2023
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Later among the works it cites.
Defending against transfer attacks from public models
Sitawarin, C., Chang, J., Huang, D., Altoyan, W., and Wagner, D · 2024
Closest in time.
Dolma: An open corpus of three trillion tokens for language model pretraining research, January 2024
Soldaini, L., Kinney, R., Bhagia, A., Schwenk, D., Atkinson, D., Authur, R., Bogin, B., Chandu, K., Dumas, J., Elazar, Y., Hofmann, V., Jha, A. H., Kumar, S., Lucy, L., Lyu, X., Lambert, N., Magnusson, I., Morrison, J., Muennighoff, N., Naik, A., Nam, C., Peters, M. E., Ravichander, A., Richardson, K., Shen, Z., Strubell, E., Subramani, N., Tafjord, O., Walsh, P., Zettlemoyer, L., Smith, N. A., Hajishirzi, H., Beltagy, I., Groeneveld, D., Dodge, J., and Lo, K · 2024
Closest in time.
All in how you ask for it: Simple black-box method for jailbreak attacks, January 2024
Takemoto, K · 2024
Closest in time.
Zeng, Y., Lin, H., Zhang, J., Yang, D., Jia, R., and Shi, W · 2024
Closest in time.