Fetching the paper…
Reading the bibliography…
The rapid advancement of large language models (LLMs) necessitates effective mechanisms to ensure their responsible deployment by accurately distinguishing unsafe content from benign content.
Explaining and harnessing adversarial examples
Goodfellow, I. J., Shlens, J., and Szegedy, C · 2014
Earlier work this paper cites.
Feature cross-substitution in adversarial classification
Li, B. and Vorobeychik, Y · 2014
Earlier work this paper cites.
Improving neural machine translation models with monolingual data
Sennrich, R., Haddow, B., and Birch, A · 2016
Earlier work this paper cites.
A game-theoretic analysis of adversarial classification
Dritsoula, L., Loiseau, P., and Musacchio, J · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Madry, A · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Understanding back-translation at scale
Edunov, S., Ott, M., Auli, M., and Grangier, D · 2018
Earlier work this paper cites.
Tagged back-translation
Caswell, I., Chelba, C., and Grangier, D · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Tagged back-translation revisited: Why does it really work?
Marie, B., Rubino, R., and Fujita, A · 2020
Earlier work this paper cites.
Recent advances in adversarial training for adversarial robustness
Bai, T., Luo, J., Zhao, J., Wen, B., and Wang, Q · 2021
Earlier work this paper cites.
Data augmentation for text generation without any augmented data
Bi, W., Li, H., and Huang, J · 2021
Earlier work this paper cites.
Back-translation for large-scale multilingual machine translation
Liao, B., Khadivi, S., and Hewavitharana, S · 2021
Earlier work this paper cites.
Adversarial machine learning in image classification: A survey toward the defender’s perspective
Machado, G. R., Silva, E., and Goldschmidt, R. R · 2021
Earlier work this paper cites.
Meta back-translation
Pham, H., Wang, X., Yang, Y., and Neubig, G · 2021
Earlier work this paper cites.
Paradetox: Detoxification with parallel data
Logacheva, V., Dementieva, D., Ustyantsev, S., Moskovskiy, D., Dale, D., Krotova, I., Semenov, N., and Panchenko, A · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Multilingual hatecheck: Functional tests for multilingual hate speech detection models
Röttger, P., Seelawi, H., Nozza, D., Talat, Z., and Vidgen, B · 2022
Earlier work this paper cites.
Adversarial classification: Necessary conditions and geometric flows
Trillos, N. G. and Murray, R · 2022
Earlier work this paper cites.
On synthetic data for back translation
Xu, J., Ruan, Y., Bi, W., Huang, G., Shi, S., Chen, L., and Liu, L · 2022
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., et al · 2023
Earlier work this paper cites.
Openorca: An open dataset of gpt augmented flan reasoning traces
Lian, W., Goodson, B., Pentland, E., Cook, A., Vong, C., and ”Teknium” · 2023
Earlier work this paper cites.
Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation
Lin, Z., Wang, Z., Tong, Y., Wang, Y., Guo, Y., Wang, Y., and Shang, J · 2023
Cited alongside, same era.
A holistic approach to undesired content detection in the real world
Markov, T., Zhang, C., Agarwal, S., Nekoul, F. E., Lee, T., Adler, S., Jiang, A., and Weng, L · 2023
Cited alongside, same era.
Nash learning from human feedback
Munos, R., Valko, M., Calandriello, D., Azar, M. G., Rowland, M., Guo, Z. D., Tang, Y., Geist, M., Mesnard, T., Michi, A., et al · 2023
Cited alongside, same era.
Gpt-4 technical report, 2023
OpenAI · 2023
Cited alongside, same era.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Qi, X., Zeng, Y., Xie, T., Chen, P.-Y., Jia, R., Mittal, P., and Henderson, P · 2023
Overview of the multilingual text detoxification task at pan 2024
Dementieva, D., Moskovskiy, D., Babakov, N., Ayele, A. A., Rizwan, N., Schneider, F., Wang, X., Yimam, S. M., Ustalov, D., Stakovskii, E., Smirnova, A., Elnagar, A., Mukherjee, A., and Panchenko, A · 2024
Later among the works it cites.
Multilingual jailbreak challenges in large language models
Deng, Y., Zhang, W., Pan, S. J., and Bing, L · 2024
Later among the works it cites.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Later among the works it cites.
Aegis: Online adaptive ai content safety moderation with ensemble of llm experts
Ghosh, S., Varshney, P., Galinkin, E., and Parisien, C · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Cited alongside, same era.
Xstest: A test suite for identifying exaggerated safety behaviours in large language models
Röttger, P., Kirk, H. R., Vidgen, B., Attanasio, G., Bianchi, F., and Hovy, D · 2023
Cited alongside, same era.
All languages matter: On the multilingual safety of large language models
Wang, W., Tu, Z., Chen, C., Yuan, Y., Huang, J.-t., Jiao, W., and Lyu, M. R · 2023
Cited alongside, same era.
Wizardlm: Empowering large language models to follow complex instructions
Xu, C., Sun, Q., Zheng, K., Geng, X., Zhao, P., Feng, J., Tao, C., and Jiang, D · 2023
Cited alongside, same era.
Multilingual content moderation: A case study on reddit
Ye, M., Sikka, K., Atwell, K., Hassan, S., Divakaran, A., and Alikhani, M · 2023
Cited alongside, same era.
Lmsys-chat-1m: A large-scale real-world llm conversation dataset
Zheng, L., Chiang, W.-L., Sheng, Y., Li, T., Zhuang, S., Wu, Z., Zhuang, Y., Li, Z., Lin, Z., Xing, E. P., et al · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., and Fredrikson, M · 2023
Cited alongside, same era.
Hosseini, A., Yuan, X., Malkin, N., Courville, A., Sordoni, A., and Agarwal, R · 2024
Later among the works it cites.
Jain, D., Kumar, P., Gehman, S., Zhou, X., Hartvigsen, T., and Sap, M · 2024
Later among the works it cites.
Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models, 2024
Jiang, L., Rao, K., Han, S., Ettinger, A., Brahman, F., Kumar, S., Mireshghallah, N., Lu, X., Sap, M., Choi, Y., and Dziri, N · 2024
Later among the works it cites.
Rewardbench: Evaluating reward models for language modeling
Lambert, N., Pyatkin, V., Morrison, J., Miranda, L., Lin, B. Y., Chandu, K., Dziri, N., Kumar, S., Zick, T., Choi, Y., et al · 2024
Later among the works it cites.
Salad-bench: A hierarchical and comprehensive safety benchmark for large language models
Li, L., Dong, B., Wang, R., Hu, X., Zuo, W., Lin, D., Qiao, Y., and Shao, J · 2024
Later among the works it cites.
Ma, H., Hu, T., Pu, Z., Liu, B., Ai, X., Liang, Y., and Chen, M · 2024
Later among the works it cites.
Qwen2.5: A party of foundation models, September 2024
Qwen Team · 2024
Later among the works it cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Zhang, M., Li, Y., Wu, Y., and Guo, D · 2024
Later among the works it cites.
” do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models
Shen, X., Chen, Z., Backes, M., Shen, Y., and Zhang, Y · 2024
Later among the works it cites.
A minimaximalist approach to reinforcement learning from human feedback
Swamy, G., Dann, C., Kidambi, R., Wu, Z. S., and Agarwal, A · 2024
Later among the works it cites.
From languages to geographies: Towards evaluating cultural bias in hate speech datasets
Tonneau, M., Liu, D., Fraiberger, S., Schroeder, R., Hale, S., and Röttger, P · 2024
Later among the works it cites.
Jailbroken: How does llm safety training fail?
Wei, A., Haghtalab, N., and Steinhardt, J · 2024
Later among the works it cites.
Self-play preference optimization for language model alignment
Wu, Y., Sun, Z., Yuan, H., Ji, K., Yang, Y., and Gu, Q · 2024
Later among the works it cites.
Sorry-bench: Systematically evaluating large language model safety refusal behaviors, 2024
Xie, T., Qi, X., Zeng, Y., Huang, Y., Sehwag, U. M., Huang, K., He, L., Wei, B., Li, D., Sheng, Y., Jia, R., Li, B., Li, K., Chen, D., Henderson, P., and Mittal, P · 2024
Later among the works it cites.
Benchmarking llm guardrails in handling multilingual toxicity
Yang, Y., Dan, S., Roth, D., and Lee, I · 2024
Later among the works it cites.
A survey on large language model (llm) security and privacy: The good, the bad, and the ugly
Yao, Y., Duan, J., Xu, K., Cai, Y., Sun, Z., and Zhang, Y · 2024
Later among the works it cites.
Reflect-rl: Two-player online rl fine-tuning for lms
Zhou, R., Du, S. S., and Li, B · 2024
Later among the works it cites.