Fetching the paper…
Reading the bibliography…
In response to rising concerns surrounding the safety, security, and trustworthiness of Generative AI (GenAI) models, practitioners and regulators alike have pointed to AI red-teaming as a key component of their strategies for identifying and mitigating these risks.
Red teaming of advanced information assurance concepts
Wood, B. J., & Duggan, R. A. (2000) · 2000
Earlier work this paper cites.
Recipes for safety in open-domain chatbots
Xu, J., Ju, D., Li, M., Boureau, Y.-L., Weston, J., & Dinan, E. (2020) · 2010
Earlier work this paper cites.
Computational red teaming: Past, present and future
Abbass, H., Bender, A., Gaidow, S., & Whitbread, P. (2011) · 2011
Earlier work this paper cites.
The future of crowd work
Kittur, A., Nickerson, J. V., Bernstein, M., Gerber, E., Shaw, A., Zimmerman, J., Lease, M., & Horton, J. (2013) · 2013
Earlier work this paper cites.
The tsa is in the business of’security theater,’not security
Levenson, E. (2014) · 2014
Earlier work this paper cites.
The Martian
Weir, A. (2014) · 2014
Earlier work this paper cites.
Computational red teaming
Abbass, H. A. (2015) · 2015
Earlier work this paper cites.
Red Team: How to succeed by thinking like the enemy
Zenko, M. (2015) · 2015
Earlier work this paper cites.
Demographics and discussion influence views on algorithmic fairness
Pierson, E. (2017) · 2017
Earlier work this paper cites.
Making better use of the crowd: How crowdsourcing can advance machine learning research
Vaughan, J. W. (2017) · 2017
Earlier work this paper cites.
Augmenting machine learning with argumentation
Bishop, M., Gates, C., & Levitt, K. (2018) · 2018
Earlier work this paper cites.
Algorithmic impact assessments: a practical framework for public agency
Reisman, D., Schultz, J., Crawford, K., & Whittaker, M. (2018) · 2018
Earlier work this paper cites.
A qualitative exploration of perceptions of algorithmic fairness
Woodruff, A., Fox, S. E., Rousso-Schindler, S., & Warshaw, J. (2018) · 2018
Earlier work this paper cites.
Efficient elicitation approaches to estimate collective crowd answers
Chung, J. J. Y., Song, J. Y., Kutty, S., Hong, S., Kim, J., & Lasecki, W. S. (2019) · 2019
Earlier work this paper cites.
Understanding and mitigating worker biases in the crowdsourced collection of subjective judgments
Hube, C., Fetahu, B., & Gadiraju, U. (2019) · 2019
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., & Smith, N. A. (2020) · 2020
Earlier work this paper cites.
Co-designing checklists to understand organizational challenges and opportunities around fairness in ai
Madaio, M. A., Stark, L., Wortman Vaughan, J., & Wallach, H. (2020) · 2020
Earlier work this paper cites.
Closing the ai accountability gap: Defining an end-to-end framework for internal algorithmic auditing
Raji, I. D., Smart, A., White, R. N., Mitchell, M., Gebru, T., Hutchinson, B., Smith-Loud, J., Theron, D., & Barnes, P. (2020) · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al. (2021) · 2021
Earlier work this paper cites.
Multimodal datasets: misogyny, pornography, and malignant stereotypes
Birhane, A., Prabhu, V. U., & Kahembwe, E. (2021) · 2021
Earlier work this paper cites.
Goldilocks: Consistent crowdsourced scalar annotations with relative uncertainty
Chen, Q. Z., Weld, D. S., & Zhang, A. X. (2021) · 2021
Earlier work this paper cites.
Algorithmic monoculture and social welfare
Kleinberg, J., & Raghavan, M. (2021) · 2021
Earlier work this paper cites.
Nahar, N., Zhou, S., Lewis, G., & Kästner, C. (2021) · 2021
Earlier work this paper cites.
Scaling language models: Methods, analysis & insights from training gopher
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., et al. (2021) · 2021
Earlier work this paper cites.
Where responsible ai meets reality: Practitioner perspectives on enablers for shifting organizational practices
Rakova, B., Yang, J., Cramer, H., & Chowdhury, R. (2021) · 2021
Earlier work this paper cites.
Hatecheck: Functional tests for hate speech detection models
Röttger, P., Vidgen, B., Nguyen, D., Waseem, Z., Margetts, H., & Pierrehumbert, J. (2021) · 2021
Earlier work this paper cites.
Reliability testing for natural language processing systems
Tan, S., Joty, S., Baxter, K., Taeihagh, A., Bennett, G. A., & Kan, M.-Y. (2021) · 2021
Earlier work this paper cites.
Quantifying the invisible labor in crowd work
Toxtli, C., Suri, S., & Savage, S. (2021) · 2021
Earlier work this paper cites.
Bot-adversarial dialogue for safe conversational agents
Xu, J., Ju, D., Li, M., Boureau, Y.-L., Weston, J., & Dinan, E. (2021) · 2021
Earlier work this paper cites.
Power to the people? opportunities and challenges for participatory ai
Birhane, A., Isaac, W., Prabhakaran, V., Diaz, M., Elish, M. C., Gabriel, I., & Mohamed, S. (2022) · 2022
Earlier work this paper cites.
Adversarial text normalization
Bitton, J., Pavlova, M., & Evtimov, I. (2022) · 2022
Earlier work this paper cites.
Picking on the same person: Does algorithmic monoculture lead to outcome homogenization?
Bommasani, R., Creel, K. A., Kumar, A., Jurafsky, D., & Liang, P. S. (2022) · 2022
Earlier work this paper cites.
Understanding implementation challenges in machine learning documentation
Chang, J., & Custis, C. (2022) · 2022
Earlier work this paper cites.
Who audits the auditors? recommendations from a field scan of the algorithmic auditing ecosystem
Costanza-Chock, S., Harvey, E., Raji, I. D., Czernuszenko, M., & Buolamwini, J. (2022) · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., et al. (2022) · 2022
Earlier work this paper cites.
On the horizon: Interactive and compositional deepfakes
Horvitz, E. (2022) · 2022
Earlier work this paper cites.
Poster cti4ai: Threat intelligence generation and sharing after red teaming ai models
Nguyen, C., Morgan, C., & Mittal, S. (2022) · 2022
Earlier work this paper cites.
Red teaming language models with language models
Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., & Irving, G. (2022) · 2022
Earlier work this paper cites.
Hierarchical text-conditional image generation with clip latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., & Chen, M. (2022) · 2022
Earlier work this paper cites.
Red-teaming the stable diffusion safety filter
Rando, J., Paleka, D., Lindner, D., Heim, L., & Tramer, F. (2022) · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., & Ommer, B. (2022) · 2022
Earlier work this paper cites.
Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection
Abdelnabi, S., Greshake, K., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023) · 2023
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. (2023) · 2023
Earlier work this paper cites.
Musiclm: Generating music from text
Agostinelli, A., Denk, T. I., Borsos, Z., Engel, J., Verzetti, M., Caillon, A., Huang, Q., Jansen, A., Roberts, A., Tagliasacchi, M., et al. (2023) · 2023
Earlier work this paper cites.
Detecting language model attacks with perplexity
Alon, G., & Kamfonas, M. (2023) · 2023
Earlier work this paper cites.
Towards publicly accountable frontier llms
Anderljung, M., Smith, E., O’Brien, J., Soder, L., Bucknall, B., Bluemke, E., Schuett, J., Trager, R., Strahm, L., & Chowdhury, R. (2023) · 2023
Earlier work this paper cites.
Frontier threats red teaming for ai safety. URL https://www.anthropic.com/news/frontier-threats-red-teaming-for-ai-safety
Anthropic (2023) · 2023
Earlier work this paper cites.
Model card and evaluations for claude models. URL https://www-cdn.anthropic.com/bd2a28d2535bfb0494cc8e2a3bf135d2e7523226/Model-Card-Claude-2.pdf
Anthropic (2023) · 2023
Earlier work this paper cites.
Dices dataset: diversity in conversational ai evaluation for safety
Aroyo, L., Taylor, A. S., Díaz, M., Homan, C. M., Parrish, A., Serapio-García, G., Prabhakaran, V., & Wang, D. (2023) · 2023
Earlier work this paper cites.
Image hijacking: Adversarial images can control generative models at runtime
Bailey, L., Ong, E., Russell, S., & Emmons, S. (2023) · 2023
Earlier work this paper cites.
Representation in ai evaluations
Bergman, A. S., Hendricks, L. A., Rauh, M., Wu, B., Agnew, W., Kunesch, M., Duan, I., Gabriel, I., & Isaac, W. (2023) · 2023
Earlier work this paper cites.
Living guidelines for generative ai — why scientists must oversee its use
Bockting, C. L., van Dis, E. A. M., van Rooij, R., Zuidema, W., & Bollen, J. (2023) · 2023
Cited alongside, same era.
Ai village at def con announces largest-ever public generative ai red team. URL https://aivillage.org/generative%20red%20team/generative-red-team/
Cattell, S., Carson, A., & Chowdhury, R. (2023) · 2023
Cited alongside, same era.
A survey on evaluation of large language models
Chang, Y., Wang, X., Wang, J., Wu, Y., Zhu, K., Chen, H., Yang, L., Yi, X., Wang, C., Wang, Y., et al. (2023) · 2023
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., & Wong, E. (2023) · 2023
Cited alongside, same era.
Prompting4debugging: Red-teaming text-to-image diffusion models by finding problematic prompts
Red-teaming large language models. URL https://huggingface.co/blog/red-teaming
Rajani, N., Lambert, N., & Tunstall, L. (2023) · 2023
Later among the works it cites.
Universal jailbreak backdoors from poisoned human feedback
Rando, J., & Tramèr, F. (2023) · 2023
Later among the works it cites.
Tricking llms into disobedience: Understanding, analyzing, and preventing jailbreaks
Rao, A., Vashistha, S., Naik, A., Aditya, S., & Choudhury, M. (2023) · 2023
Later among the works it cites.
Supporting human-ai collaboration in auditing llms with llms
Rastogi, C., Tulio Ribeiro, M., King, N., Nori, H., & Amershi, S. (2023) · 2023
Later among the works it cites.
Smoothllm: Defending large language models against jailbreaking attacks
Robey, A., Wong, E., Hassani, H., & Pappas, G. (2023) · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chin, Z.-Y., Jiang, C.-M., Huang, C.-C., Chen, P.-Y., & Chiu, W.-C. (2023) · 2023
Cited alongside, same era.
The participatory turn in ai design: Theoretical foundations and the current state of practice
Delgado, F., Yang, S., Madaio, M., & Yang, Q. (2023) · 2023
Cited alongside, same era.
Attack prompt generation for red teaming and defending large language models
Deng, B., Wang, W., Feng, F., Deng, Y., Wang, Q., & He, X. (2023a) · 2023
Cited alongside, same era.
Understanding practices, challenges, and opportunities for user-engaged algorithm auditing in industry practice
Deng, W. H., Guo, B., Devrio, A., Shen, H., Eslami, M., & Holstein, K. (2023c) · 2023
Cited alongside, same era.
Investigating practices and opportunities for cross-functional collaboration around ai fairness in industry practice
Deng, W. H., Yildirim, N., Chang, M., Eslami, M., Holstein, K., & Madaio, M. (2023d) · 2023
Cited alongside, same era.
Ding, P., Kuang, J., Ma, D., Cao, X., Xian, Y., Chen, J., & Huang, S. (2023) · 2023
Cited alongside, same era.
Singsong: Generating musical accompaniments from singing
Donahue, C., Caillon, A., Roberts, A., Manilow, E., Esling, P., Agostinelli, A., Verzetti, M., Simon, I., Pietquin, O., Zeghidour, N., et al. (2023) · 2023
Cited alongside, same era.
Analyzing the inherent response tendency of llms: Real-world instructions-driven jailbreak
Du, Y., Zhao, S., Ma, M., Chen, Y., & Qin, B. (2023) · 2023
Cited alongside, same era.
Xstest: A test suite for identifying exaggerated safety behaviours in large language models
Röttger, P., Kirk, H. R., Vidgen, B., Attanasio, G., Bianchi, F., & Hovy, D. (2023) · 2023
Later among the works it cites.
Probing llms for hate speech detection: strengths and vulnerabilities
Roy, S., Harshvardhan, A., Mukherjee, A., & Saha, P. (2023a) · 2023
Later among the works it cites.
Maatphor: Automated variant analysis for prompt injection attacks
Salem, A., Paverd, A., & Köpf, B. (2023) · 2023
Later among the works it cites.
On the adversarial robustness of multi-modal foundation models
Schlarmann, C., & Hein, M. (2023) · 2023
Later among the works it cites.
Towards best practices in agi safety and governance: A survey of expert opinion
Schuett, J., Dreksler, N., Anderljung, M., McCaffary, D., Heim, L., Bluemke, E., & Garfinkel, B. (2023) · 2023
Later among the works it cites.
Ignore this title and hackaprompt: Exposing systemic vulnerabilities of llms through a global prompt hacking competition
Schulhoff, S. V., Pinto, J., Khan, A., Bouchard, L.-F., Si, C., Anati, S., Tagliabue, V., Kost, A. L., Carnahan, C. R., & Boyd-Graber, J. L. (2023) · 2023
Later among the works it cites.
Scalable and transferable black-box jailbreaks for language models via persona modulation
Shah, R., Montixi, Q. F., Pour, S., Tagade, A., & Rando, J. (2023) · 2023
Later among the works it cites.
Shen, X., Chen, Z., Backes, M., Shen, Y., & Zhang, Y. (2023) · 2023
Later among the works it cites.
Model evaluation for extreme risks
Shevlane, T., Farquhar, S., Garfinkel, B., Phuong, M., Whittlestone, J., Leung, J., Kokotajlo, D., Marchal, N., Anderljung, M., Kolt, N., et al. (2023) · 2023
Later among the works it cites.
Red teaming language model detectors with language models
Shi, Z., Wang, Y., Yin, F., Chen, X., Chang, K.-W., & Hsieh, C.-J. (2023) · 2023
Later among the works it cites.
The gradient of generative ai release: Methods and considerations
Solaiman, I. (2023) · 2023
Later among the works it cites.
No offense taken: Eliciting offensiveness from language models
Srivastava, A., Ahuja, R., & Mukku, R. (2023) · 2023
Later among the works it cites.
Principle-driven self-alignment of language models from scratch with minimal human supervision
Sun, Z., Shen, Y., Zhou, Q., Zhang, H., Chen, Z., Cox, D., Yang, Y., & Gan, C. (2023) · 2023
Later among the works it cites.
Frontier model forum: What is red-teaming? URL https://www.frontiermodelforum.org/uploads/2023/10/FMF-AI-Red-Teaming.pdf
The Frontier Model Forum (FMF) (2023) · 2023
Later among the works it cites.
Executive order on the safe, secure, and trustworthy development and use of artificial intelligence. URL https://www.whitehouse.gov/briefing-room/presidential-actions/2023/10/30/executive-order-on-the-safe-secure-and-trustworthy-development-and-use-of-artificial-intelligence/
The White House (2023) · 2023
Later among the works it cites.
Evil geniuses: Delving into the safety of llm-based agents
Tian, Y., Yang, X., Zhang, J., Dong, Y., & Su, H. (2023) · 2023
Later among the works it cites.
Mass-producing failures of multimodal systems with language models
Tong, S., Jones, E., & Steinhardt, J. (2023) · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. (2023) · 2023
Later among the works it cites.
Ring-a-bell! how reliable are concept removal methods for diffusion models?
Tsai, Y.-L., Hsu, C.-Y., Xie, C., Lin, C.-H., Chen, J.-Y., Li, B., Chen, P.-Y., Yu, C.-M., & Huang, C.-Y. (2023) · 2023
Later among the works it cites.
How many unicorns are in this image? a safety evaluation benchmark for vision llms
Tu, H., Cui, C., Wang, Z., Zhou, Y., Zhao, B., Han, J., Zhou, W., Yao, H., & Xie, C. (2023) · 2023
Later among the works it cites.
Wang, H., & Shu, K. (2023) · 2023
Later among the works it cites.
Sociotechnical safety evaluation of generative ai systems
Weidinger, L., Rauh, M., Marchal, N., Manzini, A., Hendricks, L. A., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., et al. (2023) · 2023
Later among the works it cites.
Open (for business): Big tech, concentrated power, and the political economy of open ai
Widder, D. G., West, S., & Whittaker, M. (2023) · 2023
Later among the works it cites.
Defending chatgpt against jailbreak attack via self-reminders
Xie, Y., Yi, J., Shao, J., Curl, J., Lyu, L., Chen, Q., Xie, X., & Wu, F. (2023) · 2023
Later among the works it cites.
Cognitive overload: Jailbreaking large language models with overloaded logical thinking
Xu, N., Wang, F., Zhou, B., Li, B. Z., Xiao, C., & Chen, M. (2023) · 2023
Later among the works it cites.
Low-resource languages jailbreak gpt-4
Yong, Z. X., Menghini, C., & Bach, S. (2023) · 2023
Later among the works it cites.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts
Yu, J., Lin, X., & Xing, X. (2023) · 2023
Later among the works it cites.
Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher
Yuan, Y., Jiao, W., Wang, W., Huang, J.-t., He, P., Shi, S., & Tu, Z. (2023) · 2023
Later among the works it cites.
Trojansql: Sql injection against natural language interface to database
Zhang, J., Zhou, Y., Hui, B., Liu, Y., Li, Z., & Hu, S. (2023a) · 2023
Later among the works it cites.
Causality analysis for evaluating the security of large language models
Zhao, W., Li, Z., & Sun, J. (2023) · 2023
Later among the works it cites.
Red teaming chatgpt via jailbreaking: Bias, robustness, reliability and toxicity
Zhuo, T. Y., Huang, Y., Chen, C., & Xing, Z. (2023) · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Kolter, J. Z., & Fredrikson, M. (2023) · 2023
Later among the works it cites.
Video generation models as world simulators. URL https://openai.com/research/video-generation-models-as-world-simulators
Brooks, T., Peebles, B., Holmes, C., DePue, W., Guo, Y., Jing, L., Schnurr, D., Taylor, J., Luhman, T., Luhman, E., Ng, C., Wang, R., & Ramesh, A. (2024) · 2024
Closest in time.
Generative ai’s environmental costs are soaring — and mostly secret
Crawford, K. (2024) · 2024
Closest in time.
Openai dissolves team focused on long-term ai risks, less than one year after announcing it. URL https://www.cnbc.com/2024/05/17/openai-superalignment-sutskever-leike.html
Field, H. (2024a) · 2024
Closest in time.
Openai quietly removes ban on military use of its ai tools. URL https://www.cnbc.com/2024/01/16/openai-quietly-removes-ban-on-military-use-of-its-ai-tools.html
Field, H. (2024b) · 2024
Closest in time.
What’s in a name? auditing large language models for race and gender bias
Haim, A., Salinas, A., & Nyarko, J. (2024) · 2024
Closest in time.
Dialect prejudice predicts ai decisions about people’s character, employability, and criminality
Hofmann, V., Kalluri, P. R., Jurafsky, D., & King, S. (2024) · 2024
Closest in time.
A mechanistic understanding of alignment algorithms: A case study on dpo and toxicity
Lee, A., Bai, X., Pres, I., Wattenberg, M., Kummerfeld, J. K., & Mihalcea, R. (2024) · 2024
Closest in time.
Copilot in bing: Our approach to responsible ai. URL https://support.microsoft.com/en-us/topic/copilot-in-bing-our-approach-to-responsible-ai-45b5eae8-7466-43e1-ae98-b48f8ff8fd44
Microsoft (2024) · 2024
Closest in time.
Ai’s energy demands are out of control. welcome to the internet’s hyper-consumption era
Rogers, R. (2024) · 2024
Closest in time.
Openai working with u.s. military on cybersecurity tools. URL https://time.com/6556827/openai-us-military-cybersecurity/
Stone, B., & Bergen, M. (2024) · 2024
Closest in time.
Wan, Y., & Chang, K.-W. (2024) · 2024
Closest in time.
Sneakyprompt: Jailbreaking text-to-image generative models
Yang, Y., Hui, B., Yuan, H., Gong, N., & Cao, Y. (2024) · 2024
Closest in time.
Generative red team recap. URL https://aivillage.org/defcon%2031/generative-recap/
Cattell, S. (2023) · 2031
Closest in time.