Fetching the paper…
Reading the bibliography…
As large language models (LLMs) become increasingly capable, security and safety evaluation are crucial.
On evolution, search, optimization, genetic algorithms and martial arts : Towards memetic algorithms
Moscato, P · 1989
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., et al · 2022
Earlier work this paper cites.
Introducing ChatGPT
OpenAI · 2022
Earlier work this paper cites.
Red teaming language models with language models
Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Frontier ai regulation: Managing emerging risks to public safety
Anderljung, M., Barnhart, J., Leung, J., Korinek, A., O’Keefe, C., Whittlestone, J., Avin, S., Brundage, M., Bullock, J., Cass-Beggs, D., et al · 2023
Earlier work this paper cites.
Introducing Claude
Anthropic · 2023
Earlier work this paper cites.
Managing ai risks in an era of rapid progress
Bengio, Y., Hinton, G., Yao, A., Song, D., Abbeel, P., Harari, Y. N., Zhang, Y.-Q., Xue, L., Shalev-Shwartz, S., Hadfield, G., et al · 2023
Earlier work this paper cites.
Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence, 2023
Biden, J · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E · 2023
Earlier work this paper cites.
Generative ai in health care and liability risks for physicians and safety concerns for patients
Duffourc, M. and Gerke, S · 2023
Earlier work this paper cites.
Mart: Improving llm safety with multi-round automatic red-teaming, 2023
Ge, S., Zhou, C., Hou, R., Khabsa, M., Wang, Y.-C., Wang, Q., Han, J., and Mao, Y · 2023
Earlier work this paper cites.
Gemini: A family of highly capable multimodal models
Gemini Team · 2023
Earlier work this paper cites.
Open sesame! universal black box jailbreaking of large language models
Lapid, R., Langberg, R., and Sipper, M · 2023
Earlier work this paper cites.
Jailbreaking chatgpt via prompt engineering: An empirical study
Liu, Y., Deng, G., Xu, Z., Li, Y., Zheng, Y., Zhang, Y., Zhao, L., Zhang, T., and Liu, Y · 2023
Cited alongside, same era.
Tree of attacks: Jailbreaking black-box llms automatically
Mehrotra, A., Zampetakis, M., Kassianik, P., Nelson, B., Anderson, H., Singer, Y., and Karbasi, A · 2023
Cited alongside, same era.
GPT-4V(ision) system card
OpenAI · 2023
Cited alongside, same era.
Smoothllm: Defending large language models against jailbreaking attacks
Robey, A., Wong, E., Hassani, H., and Pappas, G. J · 2023
Cited alongside, same era.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Labash, B., Gopinath, A., Narasimhan, K., and Yao, S · 2023
Cited alongside, same era.
The ai scientist: Towards fully automated open-ended scientific discovery, 2024
Lu, C., Lu, C., Lange, R. T., Foerster, J., Clune, J., and Ha, D · 2024
Later among the works it cites.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., et al · 2024
Later among the works it cites.
Gpt-4 technical report, 2024
OpenAI · 2024
Later among the works it cites.
Rainbow teaming: Open-ended generation of diverse adversarial prompts, 2024
Samvelyan, M., Raparthy, S. C., Lupu, A., Hambro, E., Markosyan, A. H., Bhatt, M., Mao, Y., Jiang, M., Parker-Holder, J., Foerster, J., Rocktäschel, T., and Raileanu, R · 2024
Later among the works it cites.
Cheaper, better, faster, stronger
Team, M. A · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sociotechnical safety evaluation of generative ai systems
Weidinger, L., Rauh, M., Marchal, N., Manzini, A., Hendricks, L. A., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., et al · 2023
Cited alongside, same era.
ReAct: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y · 2023
Cited alongside, same era.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts
Yu, J., Lin, X., and Xing, X · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Cited alongside, same era.
Does refusal training in llms generalize to the past tense?
Andriushchenko, M. and Flammarion, N · 2024
Cited alongside, same era.
The claude 3 model family: Opus, sonnet, haiku, 2024
Anthropic · 2024
Cited alongside, same era.
Jailbreakbench: An open robustness benchmark for jailbreaking large language models
Chao, P., Debenedetti, E., Robey, A., Andriushchenko, M., Croce, F., Sehwag, V., Dobriban, E., Flammarion, N., Pappas, G. J., Tramer, F., et al · 2024
Cited alongside, same era.
the Prompter, P · 2024
Later among the works it cites.
Solving olympiad geometry without human demonstrations
Trinh, T. H., Wu, Y., Le, Q. V., He, H., and Luong, T · 2024
Later among the works it cites.
Redagent: Red teaming large language models with context-aware autonomous language agent
Xu, H., Zhang, W., Wang, Z., Xiao, F., Zheng, R., Feng, Y., Ba, Z., and Ren, K · 2024
Later among the works it cites.
Swe-agent: Agent-computer interfaces enable automated software engineering
Yang, J., Jimenez, C. E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., and Press, O · 2024
Later among the works it cites.
Ai risk categorization decoded (air 2024): From government regulations to corporate policies
Zeng, Y., Klyman, K., Zhou, A., Yang, Y., Pan, M., Jia, R., Song, D., Liang, P., and Li, B · 2024
Later among the works it cites.
Air-bench 2024: A safety benchmark based on risk categories from regulations and policies, 2024c
Zeng, Y., Yang, Y., Zhou, A., Tan, J. Z., Tu, Y., Mai, Y., Klyman, K., Pan, M., Jia, R., Song, D., Liang, P., and Li, B · 2024
Later among the works it cites.
Cybench: A framework for evaluating cybersecurity capabilities and risk of language models, 2024
Zhang, A. K., Perry, N., Dulepet, R., Jones, E., Lin, J. W., Ji, J., Menders, C., Hussein, G., Liu, S., Jasper, D., Peetathawatchai, P., Glenn, A., Sivashankar, V., Zamoshchin, D., Glikbarg, L., Askaryar, D., Yang, M., Zhang, T., Alluri, R., Tran, N., Sangpisit, R., Yiorkadjis, P., Osele, K., Raghupathi, G., Boneh, D., Ho, D. E., and Liang, P · 2024
Later among the works it cites.
Ali-agent: Assessing llms’ alignment with human values via agent-based evaluation
Zheng, J., Wang, H., Zhang, A., Nguyen, T. D., Sun, J., and Chua, T.-S · 2024
Later among the works it cites.
Robust prompt optimization for defending language models against jailbreaking attacks
Zhou, A., Li, B., and Wang, H · 2024
Later among the works it cites.
Improving alignment and robustness with circuit breakers
Zou, A., Phan, L., Wang, J., Duenas, D., Lin, M., Andriushchenko, M., Wang, R., Kolter, Z., Fredrikson, M., and Hendrycks, D · 2024
Later among the works it cites.