Fetching the paper…
Reading the bibliography…
In the burgeoning field of Large Language Models (LLMs), developing a robust safety mechanism, colloquially known as "safeguards" or "guardrails", has become imperative to ensure the ethical use of LLMs within prescribed boundaries.
E. L. Trist and K. W. Bamforth, “Studies in the quality of life: Delivered by the institute of personnel management in november 1957,” Lecture Series, 1957
1957
Earlier work this paper cites.
S. Milgram, “Behavioral study of obedience.” J. abnorm. soc. psychol. , vol. 67, no. 4, p. 371, 1963
1963
Earlier work this paper cites.
——, “Obedience to authority: An experimental view.” Contemp. Sociol. , vol. 4, no. 6, p. 617, 1975
1975
Earlier work this paper cites.
T. M. Amabile, “Social psychology of creativity: A consensual assessment technique.” J. pers. soc. psychol. , vol. 43, no. 5, p. 997, 1982
1982
Earlier work this paper cites.
A. van Lamsweerde, R. Darimont, and E. Letier, “Managing conflicts in goal-driven requirements engineering,” IEEE Trans. Softw. Eng. , vol. 24, no. 11, pp. 908–926, 1998
1998
Earlier work this paper cites.
P. Ngatchou, A. Zarei, and A. El-Sharkawi, “Pareto multi objective optimization,” in Proc. 13th Int. Conf. Intell. Syst. Appl. Power Syst. , 2005, pp. 84–91
2005
Earlier work this paper cites.
B. F. Crabtree, W. L. Miller, and K. C. Stange, “The chronic care model and diabetes management in US primary care settings: A systematic review,” Diabetes Care , vol. 34, no. 4, pp. 1058–1063, 2011
2011
Earlier work this paper cites.
O. Vinyals, M. Fortunato, and N. Jaitly, “Pointer networks,” in Adv. Neural Inf. Process. Syst. 28 (NeurIPS 2015) , C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett, Eds., vol. 28. Curran Associates, Inc., 2015
2015
Earlier work this paper cites.
M. Abadi, A. Chu, I. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang, “Deep learning with differential privacy,” in Proc. 2016 ACM SIGSAC Conf. Comput. Commun. Secur. , ser. CCS ’16. New York, NY, USA: Association for Computing Machinery, 2016, pp. 308–318
2016
Earlier work this paper cites.
D. Hendrycks and K. Gimpel, “A baseline for detecting misclassified and out-of-distribution examples in neural networks,” in 4th Int. Conf. Learn. Represent. (ICLR 2016) , 2016
2016
Earlier work this paper cites.
E. Jang, S. Gu, and B. Poole, “Categorical reparameterization with gumbel-softmax,” arXiv prepr. arXiv:1611,01144 , 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, ℒ \mathcal{L} . Kaiser, and I. Polosukhin, “Attention is all you need,” in Adv. Neural Inf. Process. Syst. 30 (NeurIPS 2017) , I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, Eds., vol. 30. Curran Associates, Inc., 2017
2017
Earlier work this paper cites.
H. Hosseini, S. Kannan, B. Zhang, and R. Poovendran, “Deceiving google’s perspective api built for detecting toxic comments,” arXiv prepr. arXiv:1702,08138 , 2017
2017
Earlier work this paper cites.
SL. Brand, J. Thompson Coon, LE. Fleming, L. Carroll, A. Bethel, and K. Wyatt, “Whole-system approaches to improving the health and wellbeing of healthcare workers: A systematic review,” PLoS ONE , vol. 12, no. 12, p. e0188418, 2017
2017
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. 2019 Conf. n. Am. Chapter Assoc. Comput. Linguist.: Hum. Lang. Technol. Minneapolis, Minnesota: Association for Computational Linguistics, Jun. 2019, pp. 4171–4186
2019
Earlier work this paper cites.
B. Goodrich, V. Rao, P. J. Liu, and M. Saleh, “Assessing the factual accuracy of generated text,” in Proc. 25th ACM SIGKDD Int. Conf. Knowl. Discov. Data Min. , 2019, pp. 166–175
2019
Earlier work this paper cites.
M. Zampieri, S. Malmasi, P. Nakov, S. Rosenthal, N. Farra, and R. Kumar, “Predicting the type and target of offensive posts in social media,” in Proc. 2019 Conf. n. Am. Chapter Assoc. Comput. Linguist.: Hum. Lang. Technol. , J. Burstein, C. Doran, and T. Solorio, Eds. Minneapolis, Minnesota: Association for Computational Linguistics, Jun. 2019, pp. 1415–1420
2019
Earlier work this paper cites.
D. Kaushik, E. Hovy, and Z. Lipton, “Learning the difference that makes a difference with counterfactually-augmented data,” in 7th Int. Conf. Learn. Represent. (ICLR 2019) , 2019
2019
Earlier work this paper cites.
E. Dinan, S. Humeau, B. Chintagunta, and J. Weston, “Build it break it fix it for dialogue safety: Robustness from adversarial human attack,” in Proc. 2019 Conf. Empir. Methods Nat. Lang. Process. 9th Int. Jt. Conf. Nat. Lang. Process. (EMNLP-IJCNLP) , 2019, pp. 4537–4546
2019
Earlier work this paper cites.
J. Cohen, E. Rosenfeld, and Z. Kolter, “Certified adversarial robustness via randomized smoothing,” in 36th Int. Conf. Mach. Learn. (ICML 2019) . PMLR, 2019, pp. 1310–1320
2019
Earlier work this paper cites.
L. Song, R. Shokri, and P. Mittal, “Privacy risks of securing machine learning models against adversarial examples,” in Proc. 2019 ACM SIGSAC Conf. Comput. Commun. Secur. London, United Kingdom: Association for Computing Machinery, 2019, pp. 241–257
2019
Earlier work this paper cites.
S. Gehman, S. Gururangan, M. Sap, Y. Choi, and N. A. Smith, “Realtoxicityprompts: Evaluating neural toxic degeneration in language models,” arXiv prepr. arXiv:2009,11462 , 2020
2020
Earlier work this paper cites.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” in Proc. 34th Int. Conf. Neural Inf. Process. Syst. , ser. NIPS’20. Red Hook, NY, USA: Curran Associates Inc., 2020
2020
Earlier work this paper cites.
M. Barrantes, B. Herudek, and R. Wang, “Adversarial nli for factual correctness in text summarisation models,” arXiv prepr. arXiv:2005,11739 , 2020
2020
Earlier work this paper cites.
S. L. Blodgett, S. Barocas, H. D. III, and H. M. Wallach, “Language (technology) is power: A critical survey of "Bias" in NLP,” in Proc. 58th Annu. Meet. Assoc. Comput. Linguist. , D. Jurafsky, J. Chai, N. Schluter, and J. R. Tetreault, Eds. Association for Computational Linguistics, 2020, pp. 5454–5476
2020
Earlier work this paper cites.
X. Ma, M. Sap, H. Rashkin, and Y. Choi, “PowerTransformer: Unsupervised controllable revision for biased language correction,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , B. Webber, T. Cohn, Y. He, and Y. Liu, Eds. Online: Association for Computational Linguistics, Nov. 2020, pp. 7426–7441
2020
Earlier work this paper cites.
J. Pavlopoulos, J. Sorensen, L. Dixon, N. Thain, and I. Androutsopoulos, “Toxicity detection: Does context really matter?” in Proc. 58th Annu. Meet. Assoc. Comput. Linguist. , D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault, Eds. Online: Association for Computational Linguistics, Jul. 2020, pp. 4296–4305
2020
Earlier work this paper cites.
M. Sap, S. Gabriel, L. Qin, D. Jurafsky, N. A. Smith, and Y. Choi, “Social bias frames: Reasoning about social and power implications of language,” in Proc. 58th Annu. Meet. Assoc. Comput. Linguist. , D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault, Eds. Online: Association for Computational Linguistics, Jul. 2020, pp. 5477–5490
2020
Earlier work this paper cites.
S. Gehman, S. Gururangan, M. Sap, Y. Choi, and N. A. Smith, “RealToxicityPrompts: Evaluating neural toxic degeneration in language models,” in Find. Assoc. Comput. Linguist.: EMNLP 2020 , T. Cohn, Y. He, and Y. Liu, Eds. Online: Association for Computational Linguistics, Nov. 2020, pp. 3356–3369
2020
Earlier work this paper cites.
C. Zeng, S. Li, Q. Li, J. Hu, and J. Hu, “A survey on machine reading comprehension—tasks, evaluation metrics and benchmark datasets,” Appl. Sci. , vol. 10, no. 21, p. 7640, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
J. Welbl, A. Glaese, J. Uesato, S. Dathathri, J. Mellor, L. A. Hendricks, K. Anderson, P. Kohli, B. Coppin, and P.-S. Huang, “Challenges in detoxifying language models,” arXiv prepr. arXiv:2109,07445 , 2021
2021
Earlier work this paper cites.
L. C. Lamb, A. d’Avila Garcez, M. Gori, M. O. Prates, P. H. Avelar, and M. Y. Vardi, “Graph neural networks meet neural-symbolic computing: A survey and perspective,” in Proc. 29th Int. Jt. Conf. Artif. Intell. (IJCAI 2021) , ser. IJCAI’20, Yokohama, Yokohama, Japan, 2021
2021
Earlier work this paper cites.
F. Nan, R. Nallapati, Z. Wang, C. N. dos Santos, H. Zhu, D. Zhang, K. McKeown, and B. Xiang, “Entity-level factual consistency of abstractive text summarization,” arXiv prepr. arXiv:2102,09130 , 2021
2021
Earlier work this paper cites.
K. Shuster, S. Poff, M. Chen, D. Kiela, and J. Weston, “Retrieval augmentation reduces hallucination in conversation,” arXiv prepr. arXiv:2104,07567 , 2021
2021
Earlier work this paper cites.
A. Mishra, D. Patel, A. Vijayakumar, X. L. Li, P. Kapanipathi, and K. Talamadupula, “Looking beyond sentence-level natural language inference for question answering and text summarization,” in Proc. 2021 Conf. n. Am. Chapter Assoc. Comput. Linguist.: Hum. Lang. Technol. , 2021, pp. 1322–1336
2021
Earlier work this paper cites.
I. Garrido-Muñoz, A. Montejo-Ráez, F. Martínez-Santiago, and L. A. Ureña-López, “A survey on bias in deep NLP,” Appl. Sci. , vol. 11, no. 7, p. 3184, 2021
2021
Earlier work this paper cites.
D. Narayanan, M. Shoeybi, J. Casper, P. LeGresley, M. Patwary, V. Korthikanti, D. Vainbrand, and B. Catanzaro, “Scaling language model training to a trillion parameters using megatron,” arXiv prepr. arXiv:2104,04473v5 , 2021
2021
Earlier work this paper cites.
S. Menini, A. P. Aprosio, and S. Tonelli, “Abuse is contextual, what about NLP? The role of context in abusive language annotation and detection,” arXiv prepr. arXiv:2103,14916 , 2021
2021
Earlier work this paper cites.
J. Welbl, A. Glaese, J. Uesato, S. Dathathri, J. Mellor, L. A. Hendricks, K. Anderson, P. Kohli, B. Coppin, and P.-S. Huang, “Challenges in detoxifying language models,” in Find. Assoc. Comput. Linguist.: EMNLP 2021 , M.-F. Moens, X. Huang, L. Specia, and S. W.-t. Yih, Eds. Association for Computational Linguistics, Nov. 2021, pp. 2447–2469
2021
Earlier work this paper cites.
U. Arora, W. Huang, and H. He, “Types of out-of-distribution texts and how to detect them,” arXiv prepr. arXiv:2109,06827 , 2021
2021
Earlier work this paper cites.
C. Guo, A. Sablayrolles, H. Jégou, and D. Kiela, “Gradient-based adversarial attacks against text transformers,” arXiv prepr. arXiv:2104,13733 , 2021
2021
Earlier work this paper cites.
X. Chen, A. Salem, D. Chen, M. Backes, S. Ma, Q. Shen, Z. Wu, and Y. Zhang, “Badnl: Backdoor attacks against nlp models with semantic-preserving improvements,” in Proc. 37th Annu. Comput. Secur. Appl. Conf. , 2021, pp. 554–569
2021
Earlier work this paper cites.
J. Mohapatra, C.-Y. Ko, L. Weng, P.-Y. Chen, S. Liu, and L. Daniel, “Hidden cost of randomized smoothing,” in Proc. 24th Int. Conf. Artif. Intell. Stat. , ser. Proceedings of Machine Learning Research, A. Banerjee and K. Fukumizu, Eds., vol. 130. PMLR, 2021-04-13/2021-04-15, pp. 4033–4041
2021
Earlier work this paper cites.
E. Perez, S. Huang, F. Song, T. Cai, R. Ring, J. Aslanides, A. Glaese, N. McAleese, and G. Irving, “Red teaming language models with language models,” arXiv prepr. arXiv:2202,03286 , 2022
2022
Earlier work this paper cites.
N. Simon and C. Muise, “TattleTale: Storytelling with planning and large language models,” in ICAPS Workshop Sched. Plan. Appl. , 2022
2022
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” NeurIPS , vol. 35, pp. 27 730–27 744, 2022
2022
Earlier work this paper cites.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al. , “Chain-of-thought prompting elicits reasoning in large language models,” NeurIPS , vol. 35, pp. 24 824–24 837, 2022
2022
Earlier work this paper cites.
O. Shaikh, H. Zhang, W. Held, M. Bernstein, and D. Yang, “On second thought, let’s not think step by step! Bias and toxicity in zero-shot reasoning,” arXiv prepr. arXiv:2212,08061 , 2022
2022
Earlier work this paper cites.
R. Qian, C. Ross, J. Fernandes, E. M. Smith, D. Kiela, and A. Williams, “Perturbation augmentation for fairer NLP,” in Proc. 2022 Conf. Empir. Methods Nat. Lang. Process. (EMNLP 2022) , Y. Goldberg, Z. Kozareva, and Y. Zhang, Eds. Association for Computational Linguistics, 2022, pp. 9496–9521
2022
Earlier work this paper cites.
C. Oh, H. Won, J. So, T. Kim, Y. Kim, H. Choi, and K. Song, “Learning fair representation via distributional contrastive disentanglement,” in 28th ACM SIGKDD Conf. Knowl. Discov. Data Min. (KDD 2022) , A. Zhang and H. Rangwala, Eds. ACM, 2022, pp. 1295–1305
2022
Earlier work this paper cites.
E. L. Ungless, A. Rafferty, H. Nag, and B. Ross, “A Robust Bias Mitigation procedure based on the stereotype content model,” arXiv prepr. arXiv:2210,14552 , 2022
2022
Earlier work this paper cites.
X. Li, F. Tramer, P. Liang, and T. Hashimoto, “Large language models can be strong differentially private learners,” in 10th Int. Conf. Learn. Represent. (ICLR 2022) , 2022
2022
Earlier work this paper cites.
F. Mireshghallah, A. Backurs, H. A. Inan, L. Wutschitz, and J. Kulkarni, “Differentially private model compression,” NeurIPS , vol. 35, pp. 29 468–29 483, 2022
2022
Earlier work this paper cites.
D. Yu, S. Naik, A. Backurs, S. Gopi, H. A. Inan, G. Kamath, J. Kulkarni, Y. T. Lee, A. Manoel, L. Wutschitz, S. Yekhanin, and H. Zhang, “Differentially private fine-tuning of language models,” in 10th Int. Conf. Learn. Represent. (ICLR 2022) . OpenReview.net, 2022
2022
Earlier work this paper cites.
W. Shi, R. Shea, S. Chen, C. Zhang, R. Jia, and Z. Yu, “Just fine-tune twice: Selective differential privacy for large language models,” arXiv prepr. arXiv:2204,07667 , 2022
2022
Earlier work this paper cites.
T. T. Nguyen, T. T. Huynh, P. L. Nguyen, A. W.-C. Liew, H. Yin, and Q. V. H. Nguyen, “A survey of machine unlearning,” arXiv prepr. arXiv:2209,02299v5 , 2022
2022
Earlier work this paper cites.
R. Plant, V. Giuffrida, and D. Gkatzia, “You are what you write: Preserving privacy in the era of large language models,” arXiv prepr. arXiv:2204,09391 , 2022
2022
Earlier work this paper cites.
X. Wang, H. Wang, and D. Yang, “Measure and improve robustness in NLP models: A survey,” arXiv prepr. arXiv:2112,08313v2 , 2022
2022
Earlier work this paper cites.
N. Goyal, I. D. Kivlichan, R. Rosen, and L. Vasserman, “Is your toxicity my toxicity? Exploring the impact of rater identity on toxicity annotation,” Proc, ACM Hum,-Comput, Interact, , vol. 6, no. CSCW2, Nov. 2022
2022
Earlier work this paper cites.
L. Rosenblatt, L. Piedras, and J. Wilkins, “Critical perspectives: A benchmark revealing pitfalls in PerspectiveAPI,” in Proc. 2nd Workshop NLP Posit. Impact (NLP4PI) , L. Biester, D. Demszky, Z. Jin, M. Sachan, J. Tetreault, S. Wilson, L. Xiao, and J. Zhao, Eds. Abu Dhabi, United Arab Emirates (Hybrid): Association for Computational Linguistics, Dec. 2022, pp. 15–24
2022
Earlier work this paper cites.
D. Ganguli, L. Lovitt, J. Kernion, A. Askell, Y. Bai, S. Kadavath, B. Mann, E. Perez, N. Schiefer, K. Ndousse et al. , “Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned,” arXiv prepr. arXiv:2209,07858 , 2022
2022
Earlier work this paper cites.
Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon et al. , “Constitutional AI: Harmlessness from AI feedback,” arXiv prepr. arXiv:2212,08073 , 2022
2022
Earlier work this paper cites.
S. Kadavath, T. Conerly, A. Askell, T. Henighan, D. Drain, E. Perez, N. Schiefer, Z. Hatfield-Dodds, N. DasSarma, E. Tran-Johnson et al. , “Language models (mostly) know what they know,” arXiv prepr. arXiv:2207,05221 , 2022
2022
Earlier work this paper cites.
S. Lin, J. Hilton, and O. Evans, “Teaching models to express their uncertainty in words,” arXiv prepr. arXiv:2205,14334 , 2022
2022
Earlier work this paper cites.
Y. Xiao, P. P. Liang, U. Bhatt, W. Neiswanger, R. Salakhutdinov, and L.-P. Morency, “Uncertainty quantification with pre-trained language models: A large-scale empirical analysis,” arXiv prepr. arXiv:2210,04714 , 2022
2022
Earlier work this paper cites.
L. Kuhn, Y. Gal, and S. Farquhar, “Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation,” in 10th Int. Conf. Learn. Represent. (ICLR 2022) , 2022
2022
Earlier work this paper cites.
F. Perez and I. Ribeiro, “Ignore previous prompt: Attack techniques for language models,” in NeuIPS Workshop Mach. Learn. Saf. , 2022
2022
Earlier work this paper cites.
X. Cai, H. Xu, S. Xu, Y. Zhang et al. , “Badprompt: Backdoor attacks on continuous prompts,” NeurIPS , vol. 35, pp. 37 068–37 080, 2022
2022
Earlier work this paper cites.
A. Xiang, “Being ’seen’ vs. ’mis-seen’: Tensions between privacy and fairness in computer vision,” Harv. J. Law Technol. , vol. 36, no. 1, 2022
2022
Earlier work this paper cites.
M. Dolata, S. Feuerriegel, and G. Schwabe, “A sociotechnical view of algorithmic fairness,” Inf. Syst. J. , vol. 32, no. 4, pp. 754–818, 2022
2022
Earlier work this paper cites.
R. Schwartz, A. Vassilev, K. Greene, L. Perine, A. Burt, and P. Hall, “Towards a standard for identifying and managing bias in artificial intelligence,” Special Publication (NIST SP), Gaithersburg, MD, 2022
2022
Earlier work this paper cites.
OpenAI, “GPT-4 technical report,” arXiv e-prints 2303,08774 , 2023
2023
Cited alongside, same era.
X. Huang, W. Ruan, W. Huang, G. Jin, Y. Dong, C. Wu, S. Bensalem, R. Mu, Y. Qi, X. Zhao et al. , “A survey of safety and trustworthiness of large language models through the lens of verification and validation,” arXiv prepr. arXiv:2305,11391 , 2023
2023
Cited alongside, same era.
D. Kang, X. Li, I. Stoica, C. Guestrin, M. Zaharia, and T. Hashimoto, “Exploiting programmatic behavior of LLMs: Dual-use through standard security attacks,” arXiv prepr. arXiv:2302,05733 , 2023
2023
Cited alongside, same era.
A. Birhane, A. Kasirzadeh, D. Leslie, and S. Wachter, “Science in the age of large language models,” Nat. Rev. Phys. , vol. 5, no. 5, pp. 277–280, May 2023
2023
Cited alongside, same era.
X. Li, Z. Zhou, J. Zhu, J. Yao, T. Liu, and B. Han, “Deepinception: Hypnotize large language model to be jailbreaker,” arXiv prepr. arXiv:2311,03191 , 2023
2023
Later among the works it cites.
Z. Wei, Y. Wang, and Y. Wang, “Jailbreak and guard aligned language models with only few in-context demonstrations,” arXiv prepr. arXiv:2310,06387 , 2023
2023
Later among the works it cites.
B. Deng, W. Wang, F. Feng, Y. Deng, Q. Wang, and X. He, “Attack prompt generation for red teaming and defending large language models,” arXiv prepr. arXiv:2310,12505 , 2023
2023
Later among the works it cites.
Y. Deng, W. Zhang, S. J. Pan, and L. Bing, “Multilingual jailbreak challenges in large language models,” in 12th Int. Conf. Learn. Represent. (ICLR 2024) , 2023
2023
Later among the works it cites.
Z. X. Yong, C. Menghini, and S. Bach, “Low-resource languages jailbreak GPT-4,” in Soc. Responsible Lang. Model. Res. , 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
T. Rebedea, R. Dinu, M. Sreedhar, C. Parisien, and J. Cohen, “Nemo guardrails: A toolkit for controllable and safe llm applications with programmable rails,” arXiv prepr. arXiv:2310,10501 , 2023
2023
Cited alongside, same era.
S. Rajpal, “Guardrails AI,” 2023
2023
Cited alongside, same era.
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar et al. , “Llama: Open and efficient foundation language models,” arXiv prepr. arXiv:2302,13971 , 2023
2023
Cited alongside, same era.
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen et al. , “Palm 2 technical report,” arXiv prepr. arXiv:2305,10403 , 2023
2023
Cited alongside, same era.
J. Wei, S. Kim, H. Jung, and Y.-H. Kim, “Leveraging large language models to power chatbots for collecting user self-reported data,” arXiv prepr. arXiv:2301,05843 , 2023
2023
Cited alongside, same era.
C. Lyu, J. Xu, and L. Wang, “New trends in machine translation using large language models: Case examples with chatgpt,” arXiv prepr. arXiv:2305,01181 , 2023
2023
Cited alongside, same era.
R. Lou, K. Zhang, and W. Yin, “Is prompt all you need? no. A comprehensive and broader view of instruction learning,” arXiv prepr. arXiv:2303,10475 , 2023
2023
Cited alongside, same era.
2023
Later among the works it cites.
P. Ding, J. Kuang, D. Ma, X. Cao, Y. Xian, J. Chen, and S. Huang, “A wolf in sheep’s clothing: Generalized nested jailbreak prompts can fool large language models easily,” arXiv prepr. arXiv:2311,08268 , 2023
2023
Later among the works it cites.
P. Chao, A. Robey, E. Dobriban, H. Hassani, G. J. Pappas, and E. Wong, “Jailbreaking black box large language models in twenty queries,” arXiv prepr. arXiv:2310,08419 , 2023
2023
Later among the works it cites.
J. Yu, X. Lin, and X. Xing, “Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts,” arXiv prepr. arXiv:2309,10253 , 2023
2023
Later among the works it cites.
A. Mehrotra, M. Zampetakis, P. Kassianik, B. Nelson, H. Anderson, Y. Singer, and A. Karbasi, “Tree of attacks: Jailbreaking black-box llms automatically,” arXiv prepr. arXiv:2312,02119 , 2023
2023
Later among the works it cites.
D. Glukhov, I. Shumailov, Y. Gal, N. Papernot, and V. Papyan, “Llm censorship: A machine learning challenge or a computer security problem?” arXiv prepr. arXiv:2307,10719 , 2023
2023
Later among the works it cites.
K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection,” in Proc. 16th ACM Workshop Artif. Intell. Secur. , 2023, pp. 79–90
2023
Later among the works it cites.
Y. Liu, G. Deng, Y. Li, K. Wang, T. Zhang, Y. Liu, H. Wang, Y. Zheng, and Y. Liu, “Prompt injection attack against LLM-integrated applications,” arXiv prepr. arXiv:2306,05499 , 2023
2023
Later among the works it cites.
R. Lapid, R. Langberg, and M. Sipper, “Open sesame! universal black box jailbreaking of large language models,” arXiv prepr. arXiv:2309,01446 , 2023
2023
Later among the works it cites.
S. Jiang, X. Chen, and R. Tang, “Prompt packer: Deceiving llms through compositional instruction with hidden attacks,” arXiv prepr. arXiv:2310,10077 , 2023
2023
Later among the works it cites.
K. Pelrine, M. Taufeeque, M. Zając, E. McLean, and A. Gleave, “Exploiting novel gpt-4 apis,” arXiv prepr. arXiv:2312,14302 , 2023
2023
Later among the works it cites.
Q. Zhan, R. Fang, R. Bindu, A. Gupta, T. Hashimoto, and D. Kang, “Removing rlhf protections in gpt-4 via fine-tuning,” arXiv prepr. arXiv:2311,05553 , 2023
2023
Later among the works it cites.
F. Bianchi, M. Suzgun, G. Attanasio, P. Röttger, D. Jurafsky, T. Hashimoto, and J. Zou, “Safety-tuned llamas: Lessons from improving the safety of large language models that follow instructions,” arXiv prepr. arXiv:2309,07875 , 2023
2023
Later among the works it cites.
X. Chen, S. Tang, R. Zhu, S. Yan, L. Jin, Z. Wang, L. Su, X. Wang, and H. Tang, “The janus interface: How fine-tuning in large language models amplifies the privacy risks,” arXiv prepr. arXiv:2310,15469 , 2023
2023
Later among the works it cites.
M. A. Shah, R. Sharma, H. Dhamyal, R. Olivier, A. Shah, J. Konan, D. Alharthi, H. T. Bukhari, M. Baali, S. Deshmukh, M. Kuhlmann, B. Raj, and R. Singh, “LoFT: Local proxy fine-tuning for improving transferability of adversarial attacks against large language model,” arXiv prepr. arXiv:2310,04445v2 , 2023
2023
Later among the works it cites.
J. Shi, Y. Liu, P. Zhou, and L. Sun, “Badgpt: Exploring security vulnerabilities of chatgpt via backdoor attacks to instructgpt,” arXiv prepr. arXiv:2304,12298 , 2023
2023
Later among the works it cites.
H. Wang and K. Shu, “Backdoor activation attack: Attack large language models using activation steering for safety-alignment,” arXiv prepr. arXiv:2311,09433 , 2023
2023
Later among the works it cites.
N. Jain, A. Schwarzschild, Y. Wen, G. Somepalli, J. Kirchenbauer, P.-y. Chiang, M. Goldblum, A. Saha, J. Geiping, and T. Goldstein, “Baseline defenses for adversarial attacks against aligned language models,” arXiv prepr. arXiv:2309,00614 , 2023
2023
Later among the works it cites.
N. Maus, P. Chao, E. Wong, and J. R. Gardner, “Black box adversarial prompting for foundation models,” in 2nd Workshop New Front. Advers. Mach. Learn. , 2023
2023
Later among the works it cites.
E. Shayegani, M. A. A. Mamun, Y. Fu, P. Zaree, Y. Dong, and N. Abu-Ghazaleh, “Survey of vulnerabilities in large language models revealed by adversarial attacks,” arXiv prepr. arXiv:2310,10844 , 2023
2023
Later among the works it cites.
G. Alon and M. Kamfonas, “Detecting language model attacks with perplexity,” arXiv prepr. arXiv:2308,14132 , 2023
2023
Later among the works it cites.
A. Robey, E. Wong, H. Hassani, and G. J. Pappas, “Smoothllm: Defending large language models against jailbreaking attacks,” arXiv prepr. arXiv:2310,03684 , 2023
2023
Later among the works it cites.
A. Helbling, M. Phute, M. Hull, and D. H. Chau, “Llm self defense: By self examination, llms know they are being tricked,” arXiv prepr. arXiv:2308,07308 , 2023
2023
Later among the works it cites.
B. Cao, Y. Cao, L. Lin, and J. Chen, “Defending against alignment-breaking attacks via robustly aligned llm,” arXiv prepr. arXiv:2309,14348 , 2023
2023
Later among the works it cites.
B. Chen, A. Paliwal, and Q. Yan, “Jailbreaker in jail: Moving target defense for large language models,” in Proc. 10th ACM Workshop Mov. Target Def. , 2023, pp. 29–32
2023
Later among the works it cites.
Y. Li, F. Wei, J. Zhao, C. Zhang, and H. Zhang, “Rain: Your language models can align themselves without finetuning,” arXiv prepr. arXiv:2309,07124 , 2023
2023
Later among the works it cites.
Z. Zhang, J. Yang, P. Ke, and M. Huang, “Defending large language models against jailbreaking attacks through goal prioritization,” arXiv prepr. arXiv:2311,09096 , 2023
2023
Later among the works it cites.
Y. Xie, J. Yi, J. Shao, J. Curl, L. Lyu, Q. Chen, X. Xie, and F. Wu, “Defending chatgpt against jailbreak attack via self-reminders,” Nat. Mach. Intell. , vol. 5, no. 12, pp. 1486–1496, 2023
2023
Later among the works it cites.
S. Ge, C. Zhou, R. Hou, M. Khabsa, Y.-C. Wang, Q. Wang, J. Han, and Y. Mao, “Mart: Improving llm safety with multi-round automatic red-teaming,” arXiv prepr. arXiv:2311,07689 , 2023
2023
Later among the works it cites.
P. Röttger, H. R. Kirk, B. Vidgen, G. Attanasio, F. Bianchi, and D. Hovy, “Xstest: A test suite for identifying exaggerated safety behaviours in large language models,” arXiv prepr. arXiv:2308,01263 , 2023
2023
Later among the works it cites.
L. Chen, M. Zaharia, and J. Zou, “How is ChatGPT’s behavior changing over time?” arXiv prepr. arXiv:2307,09009 , 2023
2023
Later among the works it cites.
A. Narayanan and S. Kapoor, “Is GPT-4 getting worse over time?” AI Snake Oil , Jul. 2023
2023
Later among the works it cites.
T. Chakrabarty, P. Laban, D. Agarwal, S. Muresan, and C.-S. Wu, “Art or artifice? large language models and the false promise of creativity,” arXiv prepr. arXiv:2309,14556 , 2023
2023
Later among the works it cites.
S. Pal, M. Bhattacharya, S.-S. Lee, and C. Chakraborty, “A domain-specific next-generation large language model (LLM) or ChatGPT is required for biomedical engineering and research,” Ann. Biomed. Eng. , 2023/07/10, 2023
2023
Later among the works it cites.
H. Sun, J. Pei, M. Choi, and D. Jurgens, “Aligning with whom? large language models have gender and racial biases in subjective nlp tasks,” arXiv prepr. arXiv:2311,09730 , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
ActiveFence, “LLM safety review: Benchmarks and analysis,” 2023
2023
Later among the works it cites.
F. Filgueiras, R. Mendonca, and V. Almeida, “Governing artificial intelligence through a sociotechnical lens,” IEEE Internet Comput. , vol. 27, no. 05, pp. 49–52, Sep. 2023
2023
Later among the works it cites.
D. Mbiazi, M. Bhange, M. Babaei, I. Sheth, and P. J. Kenfack, “Survey on AI ethics: A socio-technical perspective,” arXiv prepr. arXiv:2311,17228 , 2023
2023
Later among the works it cites.
A. Oppermann, “What is the V-model in software development?” 2023
2023
Later among the works it cites.
Y. Ruan, H. Dong, A. Wang, S. Pitis, Y. Zhou, J. Ba, Y. Dubois, C. J. Maddison, and T. Hashimoto, “Identifying the risks of lm agents with an lm-emulated sandbox,” arXiv prepr. arXiv:2309,15817 , 2023
2023
Later among the works it cites.
S. Naihin, D. Atkinson, M. Green, M. Hamadi, C. Swift, D. Schonholtz, A. T. Kalai, and D. Bau, “Testing language model agents safely in the wild,” arXiv prepr. arXiv:2311,10538 , 2023
2023
Later among the works it cites.
M. Webster and J. Schmitt, “LLM hallucinations: How to detect and prevent them with CI,” CircleCI Blog , Jan. 2024
2024
Closest in time.
B. Wang, W. Chen, H. Pei, C. Xie, M. Kang, C. Zhang, C. Xu, Z. Xiong, R. Dutta, R. Schaeffer, S. T. Truong, S. Arora, M. Mazeika, D. Hendrycks, Z. Lin, Y. Cheng, S. Koyejo, D. Song, and B. Li, “DecodingTrust: A comprehensive assessment of trustworthiness in GPT models,” arXiv prepr. arXiv: 2306,11698 , 2024
2024
Closest in time.
M. A. Rahman, L. Alqahtani, A. Albooq, and A. Ainousah, “A survey on security and privacy of large multimodal deep learning models: Teaching and learning perspective,” in 21st Learn. Technol. Conf. (L&T 2024) . IEEE, 2024, pp. 13–18
2024
Closest in time.
I. H. Sarker, “LLM potentiality and awareness: A position paper from the perspective of trustworthy and responsible AI modeling,” Authorea Prepr. , 2024
2024
Closest in time.
A. Liu, L. Pan, X. Hu, S. Meng, and L. Wen, “A semantic invariant robust watermark for large language models,” in 12th Int. Conf. Learn. Represent. (ICLR 2024) , 2024
2024
Closest in time.
H. Koh, D. Kim, M. Lee, and K. Jung, “Can LLMs recognize toxicity? Structured toxicity investigation framework and semantic-based metric,” arXiv prepr. arXiv:2402,06900v2 , 2024
2024
Closest in time.
Y. Dong, R. Mu, G. Jin, Y. Qi, J. Hu, X. Zhao, J. Meng, W. Ruan, and X. Huang, “Building guardrails for large language models,” in 41st Int. Conf. Mach. Learn. (ICML 2024) . PMLR, 2024
2024
Closest in time.
S. Geisler, T. Wollschläger, MHI. Abdalla, J. Gasteiger, and S. Günnemann, “Attacking large language models with projected gradient descent,” arXiv prepr. arXiv:2402,09154 , 2024
2024
Closest in time.
N. Mangaokar, A. Hooda, J. Choi, S. Chandrashekaran, K. Fawaz, S. Jha, and A. Prakash, “PRP: Propagating universal perturbations to attack large language model guard-rails,” arXiv prepr. arXiv:2402,15911 , 2024
2024
Closest in time.
X. Guo, F. Yu, H. Zhang, L. Qin, and B. Hu, “Cold-attack: Jailbreaking llms with stealthiness and controllability,” arXiv prepr. arXiv:2402,08679 , 2024
2024
Closest in time.
A. Wei, N. Haghtalab, and J. Steinhardt, “Jailbroken: How does llm safety training fail?” NeurIPS , vol. 36, 2024
2024
Closest in time.
T. Liu, Y. Zhang, Z. Zhao, Y. Dong, G. Meng, and K. Chen, “Making them ask and answer: Jailbreaking large language models in few queries via disguise and reconstruction,” arXiv prepr. arXiv:2402,18104 , 2024
2024
Closest in time.
Y. Yuan, W. Jiao, W. Wang, J.-t. Huang, P. He, S. Shi, and Z. Tu, “GPT-4 is too smart to be safe: Stealthy chat with LLMs via cipher,” in 12th Int. Conf. Learn. Represent. (ICLR 2024) , 2024
2024
Closest in time.
H. Lv, X. Wang, Y. Zhang, C. Huang, S. Dou, J. Ye, T. Gui, Q. Zhang, and X. Huang, “CodeChameleon: Personalized encryption framework for jailbreaking large language models,” arXiv prepr. arXiv:2402,16717 , 2024
2024
Closest in time.
W. Zhou, X. Wang, L. Xiong, H. Xia, Y. Gu, M. Chai, F. Zhu, C. Huang, S. Dou, Z. Xi et al. , “EasyJailbreak: A unified framework for jailbreaking large language models,” arXiv prepr. arXiv:2403,12171 , 2024
2024
Closest in time.
H. Jin, R. Chen, A. Zhou, J. Chen, Y. Zhang, and H. Wang, “GUARD: Role-playing to generate natural-language jailbreakings to test guideline adherence of large language models,” arXiv prepr. arXiv:2402,03299 , 2024
2024
Closest in time.
X. Qi, Y. Zeng, T. Xie, P.-Y. Chen, R. Jia, P. Mittal, and P. Henderson, “Fine-tuning aligned language models compromises safety, even when users do not intend to!” in 12th Int. Conf. Learn. Represent. (ICLR 2024) , 2024
2024
Closest in time.
W. Zou, R. Geng, B. Wang, and J. Jia, “PoisonedRAG: Knowledge poisoning attacks to retrieval-augmented generation of large language models,” arXiv prepr. arXiv:2402,07867 , 2024
2024
Closest in time.
M. Shu, J. Wang, C. Zhu, J. Geiping, C. Xiao, and T. Goldstein, “On the exploitability of instruction tuning,” NeurIPS , vol. 36, 2024
2024
Closest in time.
S. Zhao, M. Jia, L. A. Tuan, and J. Wen, “Universal vulnerabilities in large language models: In-context learning backdoor attacks,” arXiv prepr. arXiv:2401,05949 , 2024
2024
Closest in time.
A. Zhou, B. Li, and H. Wang, “Robust prompt optimization for defending language models against jailbreaking attacks,” arXiv prepr. arXiv:2401,17263 , 2024
2024
Closest in time.
Z. Xu, F. Jiang, L. Niu, J. Jia, B. Y. Lin, and R. Poovendran, “SafeDecoding: Defending against jailbreak attacks via safety-aware decoding,” arXiv prepr. arXiv:2402,08983 , 2024
2024
Closest in time.
P. R. A. S. Bassi, S. S. J. Dertkigil, and A. Cavalli, “Improving deep neural network generalization and robustness to background bias via layer-wise relevance propagation optimization,” Nat. Commun. , vol. 15, no. 1, p. 291, 2024/01/04, 2024
2024
Closest in time.
X. Tang, Q. Jin, K. Zhu, T. Yuan, Y. Zhang, W. Zhou, M. Qu, Y. Zhao, J. Tang, Z. Zhang et al. , “Prioritizing safeguarding over autonomy: Risks of LLM agents for science,” arXiv prepr. arXiv:2402,04247 , 2024
2024
Closest in time.
T. Yuan, Z. He, L. Dong, Y. Wang, R. Zhao, T. Xia, L. Xu, B. Zhou, F. Li, Z. Zhang et al. , “R-judge: Benchmarking safety risk awareness for LLM agents,” arXiv prepr. arXiv:2401,10019 , 2024
2024
Closest in time.