Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) require careful safety alignment to prevent malicious outputs.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Measuring and mitigating unintended bias in text classification
Dixon, L., Li, J., Sorensen, J., Thain, N., and Vasserman, L · 2018
Earlier work this paper cites.
Racial bias in hate speech and abusive language detection datasets
Davidson, T., Bhattacharya, D., and Weber, I · 2019
Earlier work this paper cites.
The risk of racial bias in hate speech detection
Sap, M., Card, D., Gabriel, S., Choi, Y., and Smith, N. A · 2019
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Language models are few-shot learners
Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., et al · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Earlier work this paper cites.
Recipes for safety in open-domain chatbots
Xu, J., Ju, D., Li, M., Boureau, Y.-L., Weston, J., and Dinan, E · 2020
Earlier work this paper cites.
Challenges in automated debiasing for toxic language detection
Zhou, X · 2020
Earlier work this paper cites.
Just say no: Analyzing the stance of neural dialogue generation in offensive contexts
Baheti, A., Sap, M., Ritter, A., and Riedl, M · 2021
Earlier work this paper cites.
Adversarial glue: A multi-task benchmark for robustness evaluation of language models
Wang, B., Xu, C., Wang, S., Gan, Z., Cheng, Y., Gao, J., Awadallah, A. H., and Li, B · 2021
Earlier work this paper cites.
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al · 2022
Earlier work this paper cites.
A survey on in-context learning
Dong, Q., Li, L., Dai, D., Zheng, C., Wu, Z., Chang, B., Sun, X., Xu, J., and Sui, Z · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: methods, scaling behaviors, and lessons learned. arxiv, 2022
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., et al · 2022
Earlier work this paper cites.
Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Hartvigsen, T., Gabriel, S., Palangi, H., Sap, M., Ray, D., and Kamar, E · 2022
Earlier work this paper cites.
Prosocialdialog: A prosocial backbone for conversational agents
Kim, H., Yu, Y., Jiang, L., Lu, X., Khashabi, D., Kim, G., Choi, Y., and Sap, M · 2022
Earlier work this paper cites.
Safetext: A benchmark for exploring physical safety in language models
Levy, S., Allaway, E., Subbiah, M., Chilton, L., Patton, D., McKeown, K., and Wang, W. Y · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D · 2022
Earlier work this paper cites.
Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., et al · 2023
Earlier work this paper cites.
Bianchi, F., Suzgun, M., Attanasio, G., Röttger, P., Jurafsky, D., Hashimoto, T., and Zou, J · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E · 2023
Earlier work this paper cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Earlier work this paper cites.
Safe rlhf: Safe reinforcement learning from human feedback
Dai, J., Pan, X., Sun, R., Ji, J., Xu, X., Liu, M., Wang, Y., and Yang, Y · 2023
Cited alongside, same era.
Mart: Improving llm safety with multi-round automatic red-teaming
Ge, S., Zhou, C., Hou, R., Khabsa, M., Wang, Y.-C., Wang, Q., Han, J., and Mao, Y · 2023
Cited alongside, same era.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., et al · 2023
Cited alongside, same era.
Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation
Lin, Z., Wang, Z., Tong, Y., Wang, Y., Guo, Y., Wang, Y., and Shang, J · 2023
Cited alongside, same era.
Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms
Han, S., Rao, K., Ettinger, A., Jiang, L., Lin, B. Y., Lambert, N., Choi, Y., and Dziri, N · 2024
Closest in time.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Ji, J., Liu, M., Dai, J., Pan, X., Zhang, C., Bian, C., Chen, B., Sun, R., Wang, Y., and Yang, Y · 2024
Closest in time.
Outfox: Llm-generated essay detection through in-context learning with adversarially generated examples
Koike, R., Kaneko, M., and Okazaki, N · 2024
Closest in time.
Refusal in LLMs is mediated by a single direction — LessWrong — lesswrong.com
less · 2024
Closest in time.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Qi, X., Zeng, Y., Xie, T., Chen, P.-Y., Jia, R., Mittal, P., and Henderson, P · 2023
Cited alongside, same era.
Qiu, H., Zhang, S., Li, A., He, H., and Lan, Z · 2023
Cited alongside, same era.
Smoothllm: Defending large language models against jailbreaking attacks
Robey, A., Wong, E., Hassani, H., and Pappas, G. J · 2023
Cited alongside, same era.
Xstest: A test suite for identifying exaggerated safety behaviours in large language models
Röttger, P., Kirk, H. R., Vidgen, B., Attanasio, G., Bianchi, F., and Hovy, D · 2023
Cited alongside, same era.
Alpaca: A strong, replicable instruction-following model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Cited alongside, same era.
Jailbreak and guard aligned language models with only few in-context demonstrations
Wei, Z., Wang, Y., and Wang, Y · 2023
Cited alongside, same era.
Defending chatgpt against jailbreak attack via self-reminders
Xie, Y., Yi, J., Shao, J., Curl, J., Lyu, L., Chen, Q., Xie, X., and Wu, F · 2023
Cited alongside, same era.
Xu, L., Zhao, K., Zhu, L., and Xue, H · 2023
Cited alongside, same era.
Llama3 · 2024
Closest in time.
Inadequacies of large language model benchmarks in the era of generative artificial intelligence
McIntosh, T. R., Susnjak, T., Liu, T., Watters, P., and Halgamuge, M. N · 2024
Closest in time.
mistralai/Mistral-7B-Instruct-v0.3 · Hugging Face — huggingface.co
mistral · 2024
Closest in time.
Mistral AI — Frontier AI in your hands — mistral.ai
Mistral · 2024
Closest in time.
Introducing the model spec — openai
OpenAI · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T., Alayrac, J.-b., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., et al · 2024
Closest in time.
Rainbow teaming: Open-ended generation of diverse adversarial prompts
Samvelyan, M., Raparthy, S. C., Lupu, A., Hambro, E., Markosyan, A. H., Bhatt, M., Mao, Y., Jiang, M., Parker-Holder, J., Foerster, J., et al · 2024
Closest in time.
Trustllm: Trustworthiness in large language models
Sun, L., Huang, Y., Wang, H., Wu, S., Zhang, Q., Gao, C., Huang, Y., Lyu, W., Zhang, Y., Li, X., et al · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Team, G., Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., Sifre, L., Rivière, M., Kale, M. S., Love, J., et al · 2024
Closest in time.
Alert: A comprehensive benchmark for assessing large language models’ safety through red teaming
Tedeschi, S., Friedrich, F., Schramowski, P., Kersting, K., Navigli, R., Nguyen, H., and Li, B · 2024
Closest in time.
Judging the judges: Evaluating alignment and vulnerabilities in llms-as-judges
Thakur, A. S., Choudhary, K., Ramayapally, V. S., Vaidyanathan, S., and Hupkes, D · 2024
Closest in time.
Towards safety and helpfulness balanced responses via controllable large language models
Tuan, Y.-L., Chen, X., Smith, E. M., Martin, L., Batra, S., Celikyilmaz, A., Wang, W. Y., and Bikel, D. M · 2024
Closest in time.
Defending llms against jailbreaking attacks via backtranslation
Wang, Y., Shi, Z., Bai, A., and Hsieh, C.-J · 2024
Closest in time.
Jailbroken: How does llm safety training fail?
Wei, A., Haghtalab, N., and Steinhardt, J · 2024
Closest in time.
Llm jailbreak attack versus defense techniques–a comprehensive study
Xu, Z., Liu, Y., Deng, G., Li, Y., and Picek, S · 2024
Closest in time.
Large language model as attributed training data generator: A tale of diversity and bias
Yu, Y., Zhuang, Y., Zhang, J., Meng, Y., Ratner, A. J., Krishna, R., Shen, J., and Zhang, C · 2024
Closest in time.
R-judge: Benchmarking safety risk awareness for llm agents
Yuan, T., He, Z., Dong, L., Wang, Y., Zhao, R., Xia, T., Xu, L., Zhou, B., Li, F., Zhang, Z., et al · 2024
Closest in time.
On prompt-driven safeguarding for large language models
Zheng, C., Yin, F., Zhou, H., Meng, F., Zhou, J., Chang, K.-W., Huang, M., and Peng, N · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2024
Closest in time.