Fetching the paper…
Reading the bibliography…
In recent years, AI red teaming has emerged as a practice for probing the safety and security of generative AI systems.
The economics of cybersecurity: Principles and policy options
Moore, T. (2010) · 2010
Earlier work this paper cites.
Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review
Meek, T., Barham, H., Beltaif, N., Kaadoor, A., & Akhter, T. (2016) · 2016
Earlier work this paper cites.
Tools and Weapons: The Promise and the Peril of the Digital Age
Smith, B., Browne, C., & Gates, B. (2019) · 2019
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., & Amodei, D. (2020) · 2020
Earlier work this paper cites.
Closing the ai accountability gap: Defining an end-to-end framework for internal algorithmic auditing
Raji, I. D., Smart, A., White, R. N., Mitchell, M., Gebru, T., Hutchinson, B., Smith-Loud, J., Theron, D., & Barnes, P. (2020) · 2020
Earlier work this paper cites.
Ai and the everything in the whole wide world benchmark
Raji, I. D., Bender, E. M., Paullada, A., Denton, E., & Hanna, A. (2021) · 2021
Earlier work this paper cites.
Ethical and social risks of harm from language models
Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P.-S., Cheng, M., Glaese, M., Balle, B., Kasirzadeh, A., Kenton, Z., Brown, S., Hawkins, W., Stepleton, T., Biles, C., Birhane, A., Haas, J., Rimell, L., Hendricks, L. A., Isaac, W., Legassick, S., Irving, G., & Gabriel, I. (2021) · 2021
Earlier work this paper cites.
"real attackers don’t compute gradients": Bridging the gap between adversarial ml research and practice
Apruzzese, G., Anderson, H. S., Dambra, S., Freeman, D., Pierazzi, F., & Roundy, K. A. (2022) · 2022
Earlier work this paper cites.
Language contamination helps explains the cross-lingual capabilities of English pretrained models
Blevins, T. & Zettlemoyer, L. (2022) · 2022
Earlier work this paper cites.
Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Hartvigsen, T., Gabriel, S., Palangi, H., Sap, M., Ray, D., & Kamar, E. (2022) · 2022
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., & Evans, O. (2022) · 2022
Earlier work this paper cites.
Microsoft responsible ai standard, v2
Microsoft (2022) · 2022
Earlier work this paper cites.
The fallacy of ai functionality
Raji, I. D., Kumar, I. E., Horowitz, A., & Selbst, A. (2022) · 2022
Earlier work this paper cites.
A survey of artificial intelligence challenges: Analyzing the definitions, relationships, and evolutions
Saghiri, A. M., Vahidipour, S. M., Jabbarpour, M. R., Sookhak, M., & Forestiero, A. (2022) · 2022
Earlier work this paper cites.
Taxonomy of risks posed by language models
Weidinger, L., Uesato, J., Rauh, M., Griffin, C., Huang, P.-S., Mellor, J., Glaese, A., Cheng, M., Balle, B., Kasirzadeh, A., Biles, C., Brown, S., Kenton, Z., Hawkins, W., Stepleton, T., Birhane, A., Hendricks, L. A., Rimell, L., Isaac, W., Haas, J., Legassick, S., Irving, G., & Gabriel, I. (2022) · 2022
Earlier work this paper cites.
Mega: Multilingual evaluation of generative ai
Ahuja, K., Diddee, H., Hada, R., Ochieng, M., Ramesh, K., Jain, P., Nambi, A., Ganu, T., Segal, S., Axmed, M., Bali, K., & Sitaram, S. (2023) · 2023
Earlier work this paper cites.
Purple llama cyberseceval: A secure coding benchmark for language models
Bhatt, M., Chennabasappa, S., Nikolaidis, C., Wan, S., Evtimov, I., Gabi, D., Song, D., Ahmad, F., Aschermann, C., Fontana, L., Frolov, S., Giri, R. P., Kapil, D., Kozyrakis, Y., LeBlanc, D., Milazzo, J., Straumann, A., Synnaeve, G., Vontimitta, V., Whitman, S., & Saxe, J. (2023) · 2023
Earlier work this paper cites.
Sociotechnical harms of algorithmic systems: Scoping a taxonomy for harm reduction
Shelby, R., Rismani, S., Henne, K., Moon, A., Rostamzadeh, N., Nicholas, P., Yilla-Akbari, N., Gallegos, J., Smart, A., Garcia, E., & Virk, G. (2023) · 2023
Cited alongside, same era.
Jailbroken: How does llm safety training fail?
Wei, A., Haghtalab, N., & Steinhardt, J. (2023) · 2023
Cited alongside, same era.
Sociotechnical safety evaluation of generative ai systems
Weidinger, L., Rauh, M., Marchal, N., Manzini, A., Hendricks, L. A., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., Gabriel, I., Rieser, V., & Isaac, W. (2023) · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023) · 2023
Cited alongside, same era.
Instruction-following evaluation for large language models
Zhou, J., Lu, T., Mishra, S., Brahma, S., Basu, S., Luan, Y., Zhou, D., & Hou, L. (2023) · 2023
Cited alongside, same era.
Trustworthy llms: a survey and guideline for evaluating large language models’ alignment
Liu, Y., Yao, Y., Ton, J.-F., Zhang, X., Guo, R., Cheng, H., Klochkov, Y., Taufiq, M. F., & Li, H. (2024) · 2024
Later among the works it cites.
Generative ai misuse: A taxonomy of tactics and insights from real-world data
Marchal, N., Xu, R., Elasmar, R., Gabriel, I., Goldberg, B., & Isaac, W. (2024) · 2024
Later among the works it cites.
Tree of attacks: Jailbreaking black-box llms automatically
Mehrotra, A., Zampetakis, M., Kassianik, P., Nelson, B., Anderson, H., Singer, Y., & Karbasi, A. (2024) · 2024
Later among the works it cites.
Pyrit: A framework for security risk identification and red teaming in generative ai system
Munoz, G. D. L., Minnich, A. J., Lutz, R., Lundeen, R., Dheekonda, R. S. R., Chikanov, N., Jagdagdorj, B.-E., Pouliot, M., Chawla, S., Maxwell, W., Bullwinkel, B., Pratt, K., de Gruyter, J., Siska, C., Bryan, P., Westerhoff, T., Kawaguchi, C., Seifert, C., Kumar, R. S. S., & Zunger, Y. (2024) · 2024
Later among the works it cites.
Learning to see but forgetting to follow: Visual instruction tuning makes llms more prone to jailbreak attacks
Pantazopoulos, G., Parekh, A., Nikandrou, M., & Suglia, A. (2024) · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., & Fredrikson, M. (2023) · 2023
Cited alongside, same era.
Ai auditing: The broken bus on the road to ai accountability
Birhane, A., Steed, R., Ojewale, V., Vecchione, B., & Raji, I. D. (2024) · 2024
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., & Wong, E. (2024) · 2024
Cited alongside, same era.
garak: A framework for security probing large language models
Derczynski, L., Galinkin, E., Martin, J., Majumdar, S., & Inie, N. (2024) · 2024
Cited alongside, same era.
Red-teaming for generative ai: Silver bullet or security theater?
Feffer, M., Sinha, A., Deng, W. H., Lipton, Z. C., & Heidari, H. (2024) · 2024
Cited alongside, same era.
Coercing llms to do and reveal (almost) anything
Geiping, J., Stein, A., Shu, M., Saifullah, K., Wen, Y., & Goldstein, T. (2024) · 2024
Cited alongside, same era.
Dioptra test platform
Glasbrenner, J., Booth, H., Manville, K., Sexton, J., Chisholm, M. A., Choy, H., Hand, A., Hodges, B., Scemama, P., Cousin, D., Trapnell, E., Trapnell, M., Huang, H., Rowe, P., & Byrne, A. (2024) · 2024
Cited alongside, same era.
Later among the works it cites.
Evaluating frontier models for dangerous capabilities
Phuong, M., Aitchison, M., Catt, E., Cogan, S., Kaskasoli, A., Krakovna, V., Lindner, D., Rahtz, M., Assael, Y., Hodkinson, S., Howard, H., Lieberum, T., Kumar, R., Raad, M. A., Webson, A., Ho, L., Lin, S., Farquhar, S., Hutter, M., Deletang, G., Ruoss, A., El-Sayed, S., Brown, S., Dragan, A., Shah, R., Dafoe, A., & Shevlane, T. (2024) · 2024
Later among the works it cites.
Safetywashing: Do ai safety benchmarks actually measure safety progress?
Ren, R., Basart, S., Khoja, A., Gatti, A., Phan, L., Yin, X., Mazeika, M., Pan, A., Mukobi, G., Kim, R. H., Fitz, S., & Hendrycks, D. (2024) · 2024
Later among the works it cites.
Great, now write an article about that: The crescendo multi-turn llm jailbreak attack
Russinovich, M., Salem, A., & Eldan, R. (2024) · 2024
Later among the works it cites.
The ai risk repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence
Slattery, P., Saeri, A., Grundy, E., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., & Thompson, N. (2024) · 2024
Later among the works it cites.
Evaluating the social impact of generative ai systems in systems and society
Solaiman, I., Talat, Z., Agnew, W., Ahmad, L., Baker, D., Blodgett, S. L., Chen, C., au2, H. D. I., Dodge, J., Duan, I., Evans, E., Friedrich, F., Ghosh, A., Gohar, U., Hooker, S., Jernite, Y., Kalluri, R., Lusoli, A., Leidinger, A., Lin, M., Lin, X., Luccioni, S., Mickel, J., Mitchell, M., Newman, J., Ovalle, A., Png, M.-T., Singh, S., Strait, A., Struppek, L., & Subramonian, A. (2024) · 2024
Later among the works it cites.
Safe superintelligence inc
Sutskever, I., Gross, D., & Levy, D. (2024) · 2024
Later among the works it cites.
Adversarial machine learning: A taxonomy and terminology of attacks and mitigations
Vassilev, A., Oprea, A., Fordyce, A., & Anderson, H. (2024) · 2024
Later among the works it cites.
Operationalizing a threat model for red-teaming large language models (llms)
Verma, A., Krishna, S., Gehrmann, S., Seshadri, M., Pradhan, A., Ault, T., Barrett, L., Rabinowitz, D., Doucette, J., & Phan, N. (2024) · 2024
Later among the works it cites.
The instruction hierarchy: Training llms to prioritize privileged instructions
Wallace, E., Xiao, K., Leike, R., Weng, L., Heidecke, J., & Beutel, A. (2024) · 2024
Later among the works it cites.
Decodingtrust: A comprehensive assessment of trustworthiness in gpt models
Wang, B., Chen, W., Pei, H., Xie, C., Kang, M., Zhang, C., Xu, C., Xiong, Z., Dutta, R., Schaeffer, R., Truong, S. T., Arora, S., Mazeika, M., Hendrycks, D., Lin, Z., Cheng, Y., Koyejo, S., Song, D., & Li, B. (2024) · 2024
Later among the works it cites.
What was your prompt? a remote keylogging attack on ai assistants
Weiss, R., Ayzenshteyn, D., Amit, G., & Mirsky, Y. (2024) · 2024
Later among the works it cites.
Fundamental limitations of alignment in large language models
Wolf, Y., Wies, N., Avnery, O., Levine, Y., & Shashua, A. (2024) · 2024
Later among the works it cites.