Fetching the paper…
Reading the bibliography…
The battle for AI leadership is on, with OpenAI in the United States and DeepSeek in China as key contenders.
Abbass, H., Bender, A., Gaidow, S., Whitbread, P.: Computational Red Teaming: Past, Present and Future. IEEE Computational Intelligence Magazine 6
2011
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: Attention is all you need. In: Proceedings of the 31st International Conference on Neural Information Processing Systems. p. 6000–6010. NIPS’17, Curran Associates Inc., Red Hook, NY, USA (2017)
2017
Earlier work this paper cites.
2022
Earlier work this paper cites.
Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., Irving, G.: Red Teaming Language Models with Language Models. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. pp. 3419–3448 (2022)
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Dominique, B., Piorkowski, D., Nagireddy, M., Baldini, I.: Prompt Templates: A Methodology for Improving Manual Red Teaming Performance. In: ACM CHI Conference on Human Factors in Computing Systems (2024)
2024
Earlier work this paper cites.
2024
Cited alongside, same era.
Feffer, M., Sinha, A., Deng, W.H., Lipton, Z.C., Heidari, H.: Red-Teaming for generative AI: Silver bullet or security theater? In: Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society. vol. 7, pp. 421–437 (2024)
2024
Cited alongside, same era.
Ge, S., Zhou, C., Hou, R., Khabsa, M., Wang, Y.C., Wang, Q., Han, J., Mao, Y.: MART: Improving LLM Safety with Multi-round Automatic Red-Teaming. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). pp. 1927–1937 (2024)
2024
Cited alongside, same era.
Mehrabi, N., Goyal, P., Dupuy, C., Hu, Q., Ghosh, S., Zemel, R., Chang, K.W., Galstyan, A., Gupta, R.: FLIRT: Feedback Loop In-context Red Teaming. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. pp. 703–718 (2024)
ALIA: Artificial Linguistic Intelligence for Administration. https://alia.gob.es/eng/ (2025)
2025
Closest in time.
Alia Kit. https://langtech-bsc.gitbook.io/alia-kit (2025)
2025
Closest in time.
Barcelona Supercomputing Centre (BSC-CNS) Website. https://www.bsc.es (2025)
2025
Closest in time.
MareNostrum 5. https://www.bsc.es/ca/marenostrum/marenostrum-5 (2025)
2025
Closest in time.
2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
Microsoft: Planning red teaming for large language models (LLMs) and their applications. https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/red-teaming (2024)
2024
Cited alongside, same era.
OpenAI: OpenAI’s Approach to External Red Teaming for AI Models and Systems. Tech. rep. (2024)
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Yu, J., Lin, X., Yu, Z., Xing, X.: LLM-Fuzzer: Scaling Assessment of Large Language Model Jailbreaks. In: 33rd USENIX Security Symposium (USENIX Security 24). pp. 4657–4674 (2024)
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Supplementary material. https://zenodo.org/records/15014968
Cited in the paper.
2025
Closest in time.
OpenAI: OpenAI o3-mini: Pushing the frontier of cost-effective reasoning. https://openai.com/index/openai-o3-mini/ (2025)
2025
Closest in time.
OpenAI: OpenAI o3-mini System Card. https://cdn.openai.com/o3-mini-system-card-feb10.pdf (2025)
2025
Closest in time.
Ugarte, M., Valle, P., Parejo, J.A., Segura, S., Arrieta, A.: ASTRAL: Automated Safety Testing of Large Language Models. The 6th ACM/IEEE International Conference on Automation of Software Test (AST 2025) (2025)
2025
Closest in time.