Fetching the paper…
Reading the bibliography…
Automated red-teaming has become a crucial approach for uncovering vulnerabilities in large language models (LLMs).
Challenges of real-world reinforcement learning, 2019
Dulac-Arnold, G., Mankowitz, D., and Hester, T · 1904
Earlier work this paper cites.
A critical analysis of vulnerability taxonomies
Bishop, M. and Bailey, D · 1996
Earlier work this paper cites.
Constrained markov decision processes
Altman, E · 1999
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S. J · 1999
Earlier work this paper cites.
Convex optimization
Boyd, S. P. and Vandenberghe, L · 2004
Earlier work this paper cites.
How complex systems fail
Allspaw, J. and Cook, R. I · 2010
Earlier work this paper cites.
Beyond heuristics: learning to classify vulnerabilities and predict exploits
Bozorgi, M., Saul, L. K., Savage, S., and Voelker, G. M · 2010
Earlier work this paper cites.
Constrained optimization and Lagrange multiplier methods
Bertsekas, D. P · 2014
Earlier work this paper cites.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Shin, T., Razeghi, Y., Logan IV, R. L., Wallace, E., and Singh, S · 2020
Earlier work this paper cites.
Evaluating the evaluation of diversity in natural language generation
Tevet, G. and Berant, J · 2020
Earlier work this paper cites.
Exploitability prediction of software vulnerabilities
Bhatt, N., Anand, A., and Yadavalli, V. S · 2021
Earlier work this paper cites.
Gradient-based adversarial attacks against text transformers, 2021
Guo, C., Sablayrolles, A., Jégou, H., and Kiela, D · 2021
Earlier work this paper cites.
Safe exploration by solving early terminated mdp, 2021
Sun, H., Xu, Z., Fang, M., Peng, Z., Guo, J., Dai, B., and Zhou, B · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback, 2022
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Earlier work this paper cites.
Red teaming language models with language models, 2022
Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G · 2022
Cited alongside, same era.
Reinforcement learning with sparse rewards using guidance from offline demonstration, 2022
Rengarajan, D., Vaidya, G., Sarvesh, A., Kalathil, D., and Shakkottai, S · 2022
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Cited alongside, same era.
Safe rlhf: Safe reinforcement learning from human feedback
Dai, J., Pan, X., Sun, R., Ji, J., Xu, X., Liu, M., Wang, Y., and Yang, Y · 2023
Cited alongside, same era.
Mart: Improving llm safety with multi-round automatic red-teaming, 2023
Cold-attack: Jailbreaking llms with stealthiness and controllability
Guo, X., Yu, F., Zhang, H., Qin, L., and Hu, B · 2024
Later among the works it cites.
Curiosity-driven red-teaming for large language models
Hong, Z.-W., Shenfeld, I., Wang, T.-H., Chuang, Y.-S., Pareja, A., Glass, J., Srivastava, A., and Agrawal, P · 2024
Later among the works it cites.
Rlaif vs. rlhf: Scaling reinforcement learning from human feedback with ai feedback, 2024
Lee, H., Phatale, S., Mansoor, H., Mesnard, T., Ferret, J., Lu, K., Bishop, C., Hall, E., Carbune, V., Rastogi, A., and Prakash, S · 2024
Later among the works it cites.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024
Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., Forsyth, D., and Hendrycks, D · 2024
Later among the works it cites.
Tree of attacks: Jailbreaking black-box llms automatically, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ge, S., Zhou, C., Hou, R., Khabsa, M., Wang, Y.-C., Wang, Q., Han, J., and Mao, Y · 2023
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Cited alongside, same era.
Summary of chatgpt-related research and perspective towards the future of large language models
Liu, Y., Han, T., Ma, S., Zhang, J., Yang, Y., Tian, J., He, H., Li, A., He, M., Liu, Z., Wu, Z., Zhao, L., Zhu, D., Li, X., Qiang, N., Shen, D., Liu, T., and Ge, B · 2023
Cited alongside, same era.
Confronting reward model overoptimization with constrained rlhf
Moskovitz, T., Singh, A. K., Strouse, D., Sandholm, T., Salakhutdinov, R., Dragan, A. D., and McAleer, S · 2023
Cited alongside, same era.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models, 2023
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
Cited alongside, same era.
Zephyr: Direct distillation of lm alignment, 2023
Tunstall, L., Beeching, E., Lambert, N., Rajani, N., Rasul, K., Belkada, Y., Huang, S., von Werra, L., Fourrier, C., Habib, N., Sarrazin, N., Sanseviero, O., Rush, A. M., and Wolf, T · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models, 2023
Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., and Fredrikson, M · 2023
Cited alongside, same era.
Mehrotra, A., Zampetakis, M., Kassianik, P., Nelson, B., Anderson, H., Singer, Y., and Karbasi, A · 2024
Later among the works it cites.
Llama guard 2 — model cards and prompt formats, 2024
Meta · 2024
Later among the works it cites.
Progressive safeguards for safe and model-agnostic reinforcement learning, 2024
Omi, N., Hasanbeig, H., Sharma, H., Rajamani, S. K., and Sen, S · 2024
Later among the works it cites.
Assessing the zero-shot capabilities of llms for action evaluation in rl, 2024
Pignatelli, E., Ferret, J., Rockäschel, T., Grefenstette, E., Paglieri, D., Coward, S., and Toni, L · 2024
Later among the works it cites.
Safety alignment should be made more than just a few tokens deep, 2024
Qi, X., Panda, A., Lyu, K., Ma, X., Roy, S., Beirami, A., Mittal, P., and Henderson, P · 2024
Later among the works it cites.
Rainbow teaming: Open-ended generation of diverse adversarial prompts, 2024
Samvelyan, M., Raparthy, S. C., Lupu, A., Hambro, E., Markosyan, A. H., Bhatt, M., Mao, Y., Jiang, M., Parker-Holder, J., Foerster, J., Rocktäschel, T., and Raileanu, R · 2024
Later among the works it cites.
Shen, X., Chen, Z., Backes, M., Shen, Y., and Zhang, Y · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size, 2024
Team, G., Riviere, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., Hussenot, L., Mesnard, T., Shahriari, B., Ramé, A., Ferret, J., Liu, P., Tafti, P., Friesen, A., Casbon, M., Ramos, S., Kumar, R., Lan, C. L., Jerome, S., Tsitsulin, A., Vieillard, N., Stanczyk, P., Girgin, S., Momchev, N., Hoffman, M., Thakoor, S., Grill, J.-B., Neyshabur, B., Bachem, O., Walton, A., Severyn, A., Parrish, A., Ahmad, A., Hutchison, A., Abdagic, A., Carl, A., Shen, A., Brock, A., Coenen, A., Laforge, A., Paterson, A., Bastian, B., Piot, B., Wu, B., Royal, B., Chen, C., Kumar, C., Perry, C., Welty, C., Choquette-Choo, C. A., Sinopalnikov, D., Weinberger, D., Vijaykumar, D., Rogozińska, D., Herbison, D., Bandy, E., Wang, E., Noland, E., Moreira, E., Senter, E., Eltyshev, E., Visin, F., Rasskin, G., Wei, G., Cameron, G., Martins, G., Hashemi, H., Klimczak-Plucińska, H., Batra, H., Dhand, H., Nardini, I., Mein, J., Zhou, J., Svensson, J., Stanway, J., Chan, J., Zhou, J. P., Carrasqueira, J., Iljazi, J., Becker, J., Fernandez, J., van Amersfoort, J., Gordon, J., Lipschultz, J., Newlan, J., yeong Ji, J., Mohamed, K., Badola, K., Black, K., Millican, K., McDonell, K., Nguyen, K., Sodhia, K., Greene, K., Sjoesund, L. L., Usui, L., Sifre, L., Heuermann, L., Lago, L., McNealus, L., Soares, L. B., Kilpatrick, L., Dixon, L., Martins, L., Reid, M., Singh, M., Iverson, M., Görner, M., Velloso, M., Wirth, M., Davidow, M., Miller, M., Rahtz, M., Watson, M., Risdal, M., Kazemi, M., Moynihan, M., Zhang, M., Kahng, M., Park, M., Rahman, M., Khatwani, M., Dao, N., Bardoliwalla, N., Devanathan, N., Dumai, N., Chauhan, N., Wahltinez, O., Botarda, P., Barnes, P., Barham, P., Michel, P., Jin, P., Georgiev, P., Culliton, P., Kuppala, P., Comanescu, R., Merhej, R., Jana, R., Rokni, R. A., Agarwal, R., Mullins, R., Saadat, S., Carthy, S. M., Cogan, S., Perrin, S., Arnold, S. M. R., Krause, S., Dai, S., Garg, S., Sheth, S., Ronstrom, S., Chan, S., Jordan, T., Yu, T., Eccles, T., Hennigan, T., Kocisky, T., Doshi, T., Jain, V., Yadav, V., Meshram, V., Dharmadhikari, V., Barkley, W., Wei, W., Ye, W., Han, W., Kwon, W., Xu, X., Shen, Z., Gong, Z., Wei, Z., Cotruta, V., Kirk, P., Rao, A., Giang, M., Peran, L., Warkentin, T., Collins, E., Barral, J., Ghahramani, Z., Hadsell, R., Sculley, D., Banks, J., Dragan, A., Petrov, S., Vinyals, O., Dean, J., Hassabis, D., Kavukcuoglu, K., Farabet, C., Buchatskaya, E., Borgeaud, S., Fiedel, N., Joulin, A., Kenealy, K., Dadashi, R., and Andreev, A · 2024
Later among the works it cites.
Jailbreak and guard aligned language models with only few in-context demonstrations, 2024
Wei, Z., Wang, Y., Li, A., Mo, Y., and Wang, Y · 2024
Later among the works it cites.
Zeng, Y., Lin, H., Zhang, J., Yang, D., Jia, R., and Shi, W · 2024
Later among the works it cites.
Diver-ct: Diversity-enhanced red teaming with relaxing constraints, 2024
Zhao, A., Xu, Q., Lin, M., Wang, S., jin Liu, Y., Zheng, Z., and Huang, G · 2024
Later among the works it cites.
Purple-teaming llms with adversarial defender training, 2024
Zhou, J., Li, K., Li, J., Kang, J., Hu, M., Wu, X., and Meng, H · 2024
Later among the works it cites.