Fetching the paper…
Reading the bibliography…
Recent advances in large language model assistants have made them indispensable, raising significant concerns over managing their safety.
Jointly Measuring Diversity and Quality in Text Generation Models
Montahaei, E.; Alihosseini, D.; and Baghshah, M. S. 2019 · 1904
Earlier work this paper cites.
Fine-Tuning Language Models from Human Preferences
Ziegler, D. M.; Stiennon, N.; Wu, J.; Brown, T. B.; Radford, A.; Amodei, D.; Christiano, P. F.; and Irving, G. 2019 · 1909
Earlier work this paper cites.
A mathematical theory of communication
Shannon, C. E. 1948 · 1948
Earlier work this paper cites.
Rational choice and the structure of the environment
Simon, H. A. 1956 · 1956
Earlier work this paper cites.
Constrained optimization and Lagrange multiplier methods, by D. P. Bertsekas, Academic Press, New York, 1982, 395 pp. Price: $65.00
Yurkiewicz, J. 1985 · 1982
Earlier work this paper cites.
The ‘awful idea of accountability’: inscribing people into the measurement of objects
Hoskin, K. 1996 · 1996
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S.; and Barto, A. G. 1998 · 1998
Earlier work this paper cites.
Bleu: a Method for Automatic Evaluation of Machine Translation
Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W. 2002 · 2002
Earlier work this paper cites.
Convex Optimization
Boyd, S. P.; and Vandenberghe, L. 2010 · 2010
Earlier work this paper cites.
Thinking Inside the Box: Controlling and Using an Oracle AI
Armstrong, S.; Sandberg, A.; and Bostrom, N. 2012 · 2012
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L. 2014 · 2014
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
García, J.; and Fernández, F. 2015 · 2015
Earlier work this paper cites.
A Diversity-Promoting Objective Function for Neural Conversation Models
Li, J.; Galley, M.; Brockett, C.; Gao, J.; and Dolan, B. 2016 · 2016
Earlier work this paper cites.
Practical Black-Box Attacks against Deep Learning Systems using Adversarial Examples
Papernot, N.; McDaniel, P. D.; Goodfellow, I. J.; Jha, S.; Celik, Z. B.; and Swami, A. 2016 · 2016
Earlier work this paper cites.
Quantilizers: A Safer Alternative to Maximizers for Limited Optimization
Taylor, J. 2016 · 2016
Earlier work this paper cites.
Constrained Policy Optimization
Achiam, J.; Held, D.; Tamar, A.; and Abbeel, P. 2017 · 2017
Earlier work this paper cites.
Deep Reinforcement Learning from Human Preferences
Christiano, P. F.; Leike, J.; Brown, T. B.; Martic, M.; Legg, S.; and Amodei, D. 2017 · 2017
Earlier work this paper cites.
Practical Black-Box Attacks against Machine Learning
Papernot, N.; McDaniel, P. D.; Goodfellow, I. J.; Jha, S.; Celik, Z. B.; and Swami, A. 2017 · 2017
Earlier work this paper cites.
Curiosity-driven Exploration by Self-supervised Prediction
Pathak, D.; Agrawal, P.; Efros, A. A.; and Darrell, T. 2017 · 2017
Earlier work this paper cites.
Proximal Policy Optimization Algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Earlier work this paper cites.
Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Clark, P.; Cowhey, I.; Etzioni, O.; Khot, T.; Sabharwal, A.; Schoenick, C.; and Tafjord, O. 2018 · 2018
Earlier work this paper cites.
HotFlip: White-Box Adversarial Examples for Text Classification
Ebrahimi, J.; Rao, A.; Lowd, D.; and Dou, D. 2018 · 2018
Earlier work this paper cites.
Trick Me If You Can: Human-in-the-Loop Generation of Adversarial Examples for Question Answering
Wallace, E.; Rodriguez, P.; Feng, S.; Yamada, I.; and Boyd-Graber, J. L. 2018 · 2018
Earlier work this paper cites.
Texygen: A Benchmarking Platform for Text Generation Models
Zhu, Y.; Lu, S.; Zheng, L.; Guo, J.; Zhang, W.; Wang, J.; and Yu, Y. 2018 · 2018
Earlier work this paper cites.
Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack
Dinan, E.; Humeau, S.; Chintagunta, B.; and Weston, J. 2019 · 2019
Earlier work this paper cites.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reimers, N.; and Gurevych, I. 2019 · 2019
Earlier work this paper cites.
Universal Adversarial Triggers for Attacking and Analyzing NLP
Wallace, E.; Feng, S.; Kandpal, N.; Gardner, M.; and Singh, S. 2019 · 2019
Earlier work this paper cites.
HellaSwag: Can a Machine Really Finish Your Sentence?
Zellers, R.; Holtzman, A.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019 · 2019
Earlier work this paper cites.
Seq2Sick: Evaluating the Robustness of Sequence-to-Sequence Models with Adversarial Examples
Cheng, M.; Yi, J.; Chen, P.; Zhang, H.; and Hsieh, C. 2020 · 2020
Earlier work this paper cites.
Adversarial NLI: A New Benchmark for Natural Language Understanding
Nie, Y.; Williams, A.; Dinan, E.; Bansal, M.; Weston, J.; and Kiela, D. 2020 · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Stiennon, N.; Ouyang, L.; Wu, J.; Ziegler, D. M.; Lowe, R.; Voss, C.; Radford, A.; Amodei, D.; and Christiano, P. F. 2020 · 2020
Cited alongside, same era.
MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers
Wang, W.; Wei, F.; Dong, L.; Bao, H.; Yang, N.; and Zhou, M. 2020 · 2020
Cited alongside, same era.
A General Language Assistant as a Laboratory for Alignment
Askell, A.; Bai, Y.; Chen, A.; Drain, D.; Ganguli, D.; Henighan, T.; Jones, A.; Joseph, N.; Mann, B.; DasSarma, N.; Elhage, N.; Hatfield-Dodds, Z.; Hernandez, D.; Kernion, J.; Ndousse, K.; Olsson, C.; Amodei, D.; Brown, T. B.; Clark, J.; McCandlish, S.; Olah, C.; and Kaplan, J. 2021 · 2021
Cited alongside, same era.
Improving Question Answering Model Robustness with Synthetic Adversarial Data Generation
Bartolo, M.; Thrush, T.; Jia, R.; Riedel, S.; Stenetorp, P.; and Kiela, D. 2021 · 2021
Stanford Alpaca: An Instruction-following Llama Model
Taori, R.; Gulrajani, I.; Zhang, T.; Dubois, Y.; Li, X.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023 · 2023
Later among the works it cites.
Avalon’s game of thoughts: Battle against deception through recursive contemplation
Wang, S.; Liu, C.; Zheng, Z.; Qi, S.; Chen, S.; Yang, Q.; Zhao, A.; Wang, C.; Song, S.; and Huang, G. 2023 · 2023
Later among the works it cites.
GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Yu, J.; Lin, X.; Yu, Z.; and Xing, X. 2023 · 2023
Later among the works it cites.
Universal and Transferable Adversarial Attacks on Aligned Language Models
Zou, A.; Wang, Z.; Kolter, J. Z.; and Fredrikson, M. 2023 · 2023
Later among the works it cites.
Many-shot Jailbreaking
Anil, C.; Durmus, E.; Sharma, M.; Benton, J.; Kundu, S.; Batson, J.; Rimsky, N.; Tong, M.; Mu, J.; Ford, D.; Mosconi, F.; Agrawal, R.; Schaeffer, R.; Bashkansky, N.; Svenningsen, S.; Lambert, M.; Radhakrishnan, A.; Denison, C. E.; Hubinger, E.; Bai, Y.; Bricken, T.; Maxwell, T.; Schiefer, N.; Sully, J.; Tamkin, A.; Lanham, T.; Nguyen, K.; Korbak, T.; Kaplan, J.; Ganguli, D.; Bowman, S. R.; Perez, E.; Grosse, R.; and Duvenaud, D. K. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Calvo-Fullana, M.; Paternain, S.; Chamon, L. F. O.; and Ribeiro, A. 2021 · 2021
Cited alongside, same era.
Behavior From the Void: Unsupervised Active Pre-Training
Liu, H.; and Abbeel, P. 2021 · 2021
Cited alongside, same era.
WinoGrande: an adversarial winograd schema challenge at scale
Sakaguchi, K.; Bras, R. L.; Bhagavatula, C.; and Choi, Y. 2021 · 2021
Cited alongside, same era.
Evaluating the Evaluation of Diversity in Natural Language Generation
Tevet, G.; and Berant, J. 2021 · 2021
Cited alongside, same era.
Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection
Vidgen, B.; Thrush, T.; Waseem, Z.; and Kiela, D. 2021 · 2021
Cited alongside, same era.
Constitutional AI: Harmlessness from AI Feedback
Bai, Y.; Kadavath, S.; Kundu, S.; Askell, A.; Kernion, J.; Jones, A.; Chen, A.; Goldie, A.; Mirhoseini, A.; McKinnon, C.; Chen, C.; Olsson, C.; Olah, C.; Hernandez, D.; Drain, D.; Ganguli, D.; Li, D.; Tran-Johnson, E.; Perez, E.; Kerr, J.; Mueller, J.; Ladish, J.; Landau, J.; Ndousse, K.; Lukosiute, K.; Lovitt, L.; Sellitto, M.; Elhage, N.; Schiefer, N.; Mercado, N.; DasSarma, N.; Lasenby, R.; Larson, R.; Ringer, S.; Johnston, S.; Kravec, S.; Showk, S. E.; Fort, S.; Lanham, T.; Telleen-Lawton, T.; Conerly, T.; Henighan, T.; Hume, T.; Bowman, S. R.; Hatfield-Dodds, Z.; Mann, B.; Amodei, D.; Joseph, N.; McCandlish, S.; Brown, T.; and Kaplan, J. 2022 · 2022
Cited alongside, same era.
The Vendi Score: A Diversity Evaluation Metric for Machine Learning
Friedman, D.; and Dieng, A. B. 2022 · 2022
Cited alongside, same era.
Closest in time.
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning
Beutel, A.; Xiao, K.; hannes Heidecke, J.; and Weng, L. 2024 · 2024
Closest in time.
Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks
Cheng, Y.; Georgopoulos, M.; Cevher, V.; and Chrysos, G. G. 2024 · 2024
Closest in time.
Breaking Down the Defenses: A Comparative Survey of Attacks on Large Language Models
Chowdhury, A. G.; Islam, M. M.; Kumar, V.; Shezan, F. H.; Kumar, V.; Jain, V.; and Chadha, A. 2024 · 2024
Closest in time.
Safe RLHF: Safe Reinforcement Learning from Human Feedback
Dai, J.; Pan, X.; Sun, R.; Ji, J.; Xu, X.; Liu, M.; Wang, Y.; and Yang, Y. 2024 · 2024
Closest in time.
Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
Dong, Z.; Zhou, Z.; Yang, C.; Shao, J.; and Qiao, Y. 2024 · 2024
Closest in time.
Scaling rectified flow transformers for high-resolution image synthesis
Esser, P.; Kulal, S.; Blattmann, A.; Entezari, R.; Müller, J.; Saini, H.; Levi, Y.; Lorenz, D.; Sauer, A.; Boesel, F.; et al. 2024 · 2024
Closest in time.
Ruby Teaming: Improving Quality Diversity Search with Memory for Automated Red Teaming
Han, V. T. Y.; Bhardwaj, R.; and Poria, S. 2024 · 2024
Closest in time.
Curiosity-driven Red-teaming for Large Language Models
Hong, Z.-W.; Shenfeld, I.; Wang, T.-H.; Chuang, Y.-S.; Pareja, A.; Glass, J. R.; Srivastava, A.; and Agrawal, P. 2024 · 2024
Closest in time.
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Hubinger, E.; Denison, C.; Mu, J.; Lambert, M.; Tong, M.; MacDiarmid, M.; Lanham, T.; Ziegler, D. M.; Maxwell, T.; Cheng, N.; Jermyn, A. S.; Askell, A.; Radhakrishnan, A.; Anil, C.; Duvenaud, D.; Ganguli, D.; Barez, F.; Clark, J.; Ndousse, K.; Sachan, K.; Sellitto, M.; Sharma, M.; DasSarma, N.; Grosse, R.; Kravec, S.; Bai, Y.; Witten, Z.; Favaro, M.; Brauner, J.; Karnofsky, H.; Christiano, P. F.; Bowman, S. R.; Graham, L.; Kaplan, J.; Mindermann, S.; Greenblatt, R.; Shlegeris, B.; Schiefer, N.; and Perez, E. 2024 · 2024
Closest in time.
Reinforcement Learning from Human Feedback with Active Queries
Ji, K.; He, J.; and Gu, Q. 2024 · 2024
Closest in time.
Understanding the Effects of RLHF on LLM Generalisation and Diversity
Kirk, R.; Mediratta, I.; Nalmpantis, C.; Luketina, J.; Hambro, E.; Grefenstette, E.; and Raileanu, R. 2024 · 2024
Closest in time.
Learning diverse attacks on large language models for robust red-teaming and safety tuning
Lee, S.; Kim, M.; Cherif, L.; Dobre, D.; Lee, J.; Hwang, S. J.; Kawaguchi, K.; Gidel, G.; Bengio, Y.; Malkin, N.; and Jain, M. 2024 · 2024
Closest in time.
LLM-based Optimization of Compound AI Systems: A Survey
Lin, M.; Sheng, J.; Zhao, A.; Wang, S.; Yue, Y.; Wu, Y.; Liu, H.; Liu, J.; Huang, G.; and Liu, Y.-J. 2024 · 2024
Closest in time.
Confronting Reward Model Overoptimization with Constrained RLHF
Moskovitz, T.; Singh, A. K.; Strouse, D.; Sandholm, T.; Salakhutdinov, R.; Dragan, A.; and McAleer, S. M. 2024 · 2024
Closest in time.
Ferret: Faster and Effective Automated Red Teaming with Reward-Based Scoring Technique
Pala, T. D.; Toh, V. Y.; Bhardwaj, R.; and Poria, S. 2024 · 2024
Closest in time.
Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts
Samvelyan, M.; Raparthy, S. C.; Lupu, A.; Hambro, E.; Markosyan, A. H.; Bhatt, M.; Mao, Y.; Jiang, M.; Parker-Holder, J.; Foerster, J.; Rocktäschel, T.; and Raileanu, R. 2024 · 2024
Closest in time.
Meta Llama Guard 2
Team, L. 2024 · 2024
Closest in time.
Gradient-Based Language Model Red Teaming
Wichers, N.; Denison, C.; and Beirami, A. 2024 · 2024
Closest in time.
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
Yi, S.; Liu, Y.; Sun, Z.; Cong, T.; He, X.; Song, J.; Xu, K.; and Li, Q. 2024 · 2024
Closest in time.
Holistic Automated Red Teaming for Large Language Models through Top-Down Test Case Generation and Multi-turn Interaction
Zhang, J.; Zhou, Y.; Liu, Y.; Li, Z.; and Hu, S. 2024 · 2024
Closest in time.
Improving Diversity of Commonsense Generation by Large Language Models via In-Context Learning
Zhang, T.; Peng, B.; and Bollegala, D. 2024 · 2024
Closest in time.
ExpeL: LLM Agents Are Experiential Learners
Zhao, A.; Huang, D.; Xu, Q.; Lin, M.; Liu, Y.; and Huang, G. 2024 · 2024
Closest in time.
Efficient diffusion transformer with step-wise dynamic attention mediators
Pu, Y.; Xia, Z.; Guo, J.; Han, D.; Li, Q.; Li, D.; Yuan, Y.; Li, J.; Han, Y.; Song, S.; et al. 2025 · 2025
Closest in time.