Fetching the paper…
Reading the bibliography…
The last two years have seen a rapid growth in concerns around the safety of large language models (LLMs).
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
RedditBias: A Real-World Resource for Bias Evaluation and Debiasing of Conversational Language Models
Barikeri, S.; Lauscher, A.; Vulić, I.; and Glavaš, G. 2021 · 1955
Earlier work this paper cites.
CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models
Nangia, N.; Vania, C.; Bhalerao, R.; and Bowman, S. R. 2020 · 1967
Earlier work this paper cites.
On achieving and evaluating language-independence in NLP
Bender, E. M. 2011 · 2011
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Bolukbasi, T.; Chang, K.-W.; Zou, J. Y.; Saligrama, V.; and Kalai, A. T. 2016 · 2016
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Caliskan, A.; Bryson, J. J.; and Narayanan, A. 2017 · 2017
Earlier work this paper cites.
Word embeddings quantify 100 years of gender and ethnic stereotypes
Garg, N.; Schiebinger, L.; Jurafsky, D.; and Zou, J. 2018 · 2018
Earlier work this paper cites.
Gender Bias in Coreference Resolution
Rudinger, R.; Naradowsky, J.; Leonard, B.; and Van Durme, B. 2018 · 2018
Earlier work this paper cites.
Gender Bias in Coreference Resolution: Evaluation and Debiasing Methods
Zhao, J.; Wang, T.; Yatskar, M.; Ordonez, V.; and Chang, K.-W. 2018 · 2018
Earlier work this paper cites.
Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack
Dinan, E.; Humeau, S.; Chintagunta, B.; and Weston, J. 2019 · 2019
Earlier work this paper cites.
DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs
Dua, D.; Wang, Y.; Dasigi, P.; Stanovsky, G.; Singh, S.; and Gardner, M. 2019 · 2019
Earlier work this paper cites.
Towards Empathetic Open-domain Conversation Models: A New Benchmark and Dataset
Rashkin, H.; Smith, E. M.; Li, M.; and Boureau, Y.-L. 2019 · 2019
Earlier work this paper cites.
For: A dataset for synthetic speech detection
Reimao, R.; and Tzerpos, V. 2019 · 2019
Earlier work this paper cites.
The Woman Worked as a Babysitter: On Biases in Language Generation
Sheng, E.; Chang, K.-W.; Natarajan, P.; and Peng, N. 2019 · 2019
Earlier work this paper cites.
Social Chemistry 101: Learning to Reason about Social and Moral Norms
Forbes, M.; Hwang, J. D.; Shwartz, V.; Sap, M.; and Choi, Y. 2020 · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Gehman, S.; Gururangan, S.; Sap, M.; Choi, Y.; and Smith, N. A. 2020 · 2020
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2020 · 2020
Earlier work this paper cites.
The State and Fate of Linguistic Diversity and Inclusion in the NLP World
Joshi, P.; Santy, S.; Budhiraja, A.; Bali, K.; and Choudhury, M. 2020 · 2020
Earlier work this paper cites.
Racial disparities in automated speech recognition
Koenecke, A.; Nam, A.; Lake, E.; Nudell, J.; Quartey, M.; Mengesha, Z.; Toups, C.; Rickford, J. R.; Jurafsky, D.; and Goel, S. 2020 · 2020
Earlier work this paper cites.
UNQOVERing Stereotyping Biases via Underspecified Questions
Li, T.; Khashabi, D.; Khot, T.; Sabharwal, A.; and Srikumar, V. 2020 · 2020
Earlier work this paper cites.
Artie Bias Corpus: An Open Dataset for Detecting Demographic Bias in Speech Applications
Meyer, J.; Rauchenstein, L.; Eisenberg, J. D.; and Howell, N. 2020 · 2020
Earlier work this paper cites.
Diagnosing gender bias in image recognition systems
Schwemmer, C.; Knight, C.; Bello-Pardo, E. D.; Oklobdzija, S.; Schoonvelde, M.; and Lockhart, J. W. 2020 · 2020
Earlier work this paper cites.
Mitigating Language-Dependent Ethnic Bias in BERT
Ahn, J.; and Oh, A. 2021 · 2021
Earlier work this paper cites.
ConvAbuse: Data, Analysis, and Benchmarks for Nuanced Detection in Conversational AI
Cercas Curry, A.; Abercrombie, G.; and Rieser, V. 2021 · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M.; Tworek, J.; Jun, H.; Yuan, Q.; Pinto, H. P. d. O.; Kaplan, J.; Edwards, H.; Burda, Y.; Joseph, N.; Brockman, G.; et al. 2021 · 2021
Earlier work this paper cites.
BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation
Dhamala, J.; Sun, T.; Kumar, V.; Krishna, S.; Pruksachatkun, Y.; Chang, K.-W.; and Gupta, R. 2021 · 2021
Earlier work this paper cites.
Moral Stories: Situated Reasoning about Norms, Intents, Actions, and their Consequences
Emelin, D.; Le Bras, R.; Hwang, J. D.; Forbes, M.; and Choi, Y. 2021 · 2021
Earlier work this paper cites.
A framework for few-shot language model evaluation
Gao, L.; Tow, J.; Biderman, S.; Black, S.; DiPofi, A.; Foster, C.; Golding, L.; Hsu, J.; McDonell, K.; Muennighoff, N.; Phang, J.; Reynolds, L.; Tang, E.; Thite, A.; Wang, B.; Wang, K.; and Zou, A. 2021 · 2021
Earlier work this paper cites.
The Swedish Winogender Dataset
Hansson, S.; Mavromatakis, K.; Adesam, Y.; Bouma, G.; and Dannélls, D. 2021 · 2021
Earlier work this paper cites.
Aligning AI With Shared Human Values
Hendrycks, D.; Burns, C.; Basart, S.; Critch, A.; Li, J.; Song, D.; and Steinhardt, J. 2021 · 2021
Earlier work this paper cites.
Bias out-of-the-box: An empirical analysis of intersectional occupational biases in popular generative language models
Kirk, H. R.; Jun, Y.; Volpin, F.; Iqbal, H.; Benussi, E.; Dreyer, F.; Shtedritski, A.; and Asano, Y. 2021 · 2021
Earlier work this paper cites.
Towards understanding and mitigating social biases in language models
Liang, P. P.; Wu, C.; Morency, L.-P.; and Salakhutdinov, R. 2021 · 2021
Earlier work this paper cites.
Scruples: A corpus of community ethical judgments on 32,000 real-life anecdotes
Lourie, N.; Le Bras, R.; and Choi, Y. 2021 · 2021
Earlier work this paper cites.
StereoSet: Measuring stereotypical bias in pretrained language models
Nadeem, M.; Bethke, A.; and Reddy, S. 2021 · 2021
Earlier work this paper cites.
HONEST: Measuring Hurtful Sentence Completion in Language Models
Nozza, D.; Bianchi, F.; and Hovy, D. 2021 · 2021
Earlier work this paper cites.
Analyzing Stereotypes in Generative Text Inference Tasks
Sotnikova, A.; Cao, Y. T.; Daumé III, H.; and Rudinger, R. 2021 · 2021
Earlier work this paper cites.
Bot-Adversarial Dialogue for Safe Conversational Agents
Xu, J.; Ju, D.; Li, M.; Boureau, Y.-L.; Weston, J.; and Dinan, E. 2021 · 2021
Earlier work this paper cites.
Understanding and evaluating racial biases in image captioning
Zhao, D.; Wang, A.; and Russakovsky, O. 2021 · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y.; Jones, A.; Ndousse, K.; Askell, A.; Chen, A.; DasSarma, N.; Drain, D.; Fort, S.; Ganguli, D.; Henighan, T.; et al. 2022 · 2022
Earlier work this paper cites.
Re-contextualizing Fairness in NLP: The Case of India
Bhatt, S.; Dev, S.; Talukdar, P.; Dave, S.; and Prabhakaran, V. 2022 · 2022
Earlier work this paper cites.
SafetyKit: First Aid for Measuring Safety in Open-domain Conversational Systems
Dinan, E.; Abercrombie, G.; Bergman, A.; Spruit, S.; Hovy, D.; Boureau, Y.-L.; and Rieser, V. 2022 · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Ganguli, D.; Lovitt, L.; Kernion, J.; Askell, A.; Bai, Y.; Kadavath, S.; Mann, B.; Perez, E.; Schiefer, N.; Ndousse, K.; et al. 2022 · 2022
Earlier work this paper cites.
ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection
Hartvigsen, T.; Gabriel, S.; Palangi, H.; Sap, M.; Ray, D.; and Kamar, E. 2022 · 2022
Earlier work this paper cites.
What Would Jiminy Cricket Do? Towards Agents That Behave Morally
Hendrycks, D.; Mazeika, M.; Zou, A.; Patel, S.; Zhu, C.; Navarro, J.; Song, D.; Li, B.; and Steinhardt, J. 2022 · 2022
Earlier work this paper cites.
When to Make Exceptions: Exploring Language Models as Accounts of Human Moral Judgment
Jin, Z.; Levine, S.; Adauto, F. G.; Kamal, O.; Sap, M.; Sachan, M.; Mihalcea, R.; Tenenbaum, J. B.; and Schölkopf, B. 2022 · 2022
Earlier work this paper cites.
ProsocialDialog: A Prosocial Backbone for Conversational Agents
Kim, H.; Yu, Y.; Jiang, L.; Lu, X.; Khashabi, D.; Kim, G.; Choi, Y.; and Sap, M. 2022 · 2022
Earlier work this paper cites.
SafeText: A Benchmark for Exploring Physical Safety in Language Models
Levy, S.; Allaway, E.; Subbiah, M.; Chilton, L.; Patton, D.; McKeown, K.; and Wang, W. Y. 2022 · 2022
Earlier work this paper cites.
TruthfulQA: Measuring How Models Mimic Human Falsehoods
Lin, S.; Hilton, J.; and Evans, O. 2022 · 2022
Earlier work this paper cites.
French CrowS-Pairs: Extending a challenge dataset for measuring social bias in masked language models to a language other than English
Névéol, A.; Dupont, Y.; Bezançon, J.; and Fort, K. 2022 · 2022
Earlier work this paper cites.
Towards the detection of diffusion model deepfakes
Ricker, J.; Damm, S.; Holz, T.; and Fischer, A. 2022 · 2022
Earlier work this paper cites.
SecurityEval dataset: mining vulnerability examples to evaluate machine learning-based code generation techniques
Siddiq, M. L.; and Santos, J. C. S. 2022 · 2022
Earlier work this paper cites.
“I’m sorry to hear that”: Finding New Biases in Language Models with a Holistic Descriptor Dataset
Smith, E. M.; Hall, M.; Kambadur, M.; Presani, E.; and Williams, A. 2022 · 2022
Cited alongside, same era.
On the Safety of Conversational Models: Taxonomy, Dataset, and Benchmark
Sun, H.; Xu, G.; Deng, J.; Cheng, J.; Zheng, C.; Zhou, H.; Peng, N.; Zhu, X.; and Huang, M. 2022 · 2022
Cited alongside, same era.
SaFeRDialogues: Taking Feedback Gracefully after Conversational Safety Failures
Ung, M.; Xu, J.; and Boureau, Y.-L. 2022 · 2022
Cited alongside, same era.
Towards Identifying Social Bias in Dialog Systems: Framework, Dataset, and Benchmark
Zhou, J.; Deng, J.; Mi, F.; Li, Y.; Wang, Y.; Huang, M.; Jiang, X.; Liu, Q.; and Meng, H. 2022 · 2022
Cited alongside, same era.
The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems
Ziems, C.; Yu, J.; Wang, Y.-C.; Halevy, A.; and Yang, D. 2022 · 2022
Cited alongside, same era.
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Chao, P.; Debenedetti, E.; Robey, A.; Andriushchenko, M.; Croce, F.; Sehwag, V.; Dobriban, E.; Flammarion, N.; Pappas, G. J.; Tramèr, F.; Hassani, H.; and Wong, E. 2024 · 2024
Closest in time.
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Chiang, W.-L.; Zheng, L.; Sheng, Y.; Angelopoulos, A. N.; Li, T.; Li, D.; Zhang, H.; Zhu, B.; Jordan, M.; Gonzalez, J. E.; and Stoica, I. 2024 · 2024
Closest in time.
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Cui, J.; Chiang, W.-L.; Stoica, I.; and Hsieh, C.-J. 2024 · 2024
Closest in time.
Gemma 2: Improving Open Language Models at a Practical Size
DeepMind, G. 2024 · 2024
Closest in time.
Towards Measuring the Representation of Subjective Global Opinions in Language Models
Durmus, E.; Nguyen, K.; Liao, T.; Schiefer, N.; Askell, A.; Bakhtin, A.; Chen, C.; Hatfield-Dodds, Z.; Hernandez, D.; Joseph, N.; Lovitt, L.; McCandlish, S.; Sikder, O.; Tamkin, A.; Thamkul, J.; Kaplan, J.; Clark, J.; and Ganguli, D. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dices dataset: Diversity in conversational ai evaluation for safety
Aroyo, L.; Taylor, A.; Diaz, M.; Homan, C.; Parrish, A.; Serapio-García, G.; Prabhakaran, V.; and Wang, D. 2023 · 2023
Cited alongside, same era.
Open LLM Leaderboard
Beeching, E.; Fourrier, C.; Habib, N.; Han, S.; Lambert, N.; Rajani, N.; Sanseviero, O.; Tunstall, L.; and Wolf, T. 2023 · 2023
Cited alongside, same era.
Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment
Bhardwaj, R.; and Poria, S. 2023 · 2023
Cited alongside, same era.
Purple llama cyberseceval: A secure coding benchmark for language models
Bhatt, M.; Chennabasappa, S.; Nikolaidis, C.; Wan, S.; Evtimov, I.; Gabi, D.; Song, D.; Ahmad, F.; Aschermann, C.; Fontana, L.; et al. 2023 · 2023
Cited alongside, same era.
Easily Accessible Text-to-Image Generation Amplifies Demographic Stereotypes at Large Scale
Bianchi, F.; Kalluri, P.; Durmus, E.; Ladhak, F.; Cheng, M.; Nozza, D.; Hashimoto, T.; Jurafsky, D.; Zou, J.; and Caliskan, A. 2023 · 2023
Cited alongside, same era.
Are aligned neural networks adversarially aligned?
Carlini, N.; Nasr, M.; Choquette-Choo, C. A.; Jagielski, M.; Gao, I.; Koh, P. W.; Ippolito, D.; Tramèr, F.; and Schmidt, L. 2023 · 2023
Cited alongside, same era.
Can Large Language Models Provide Security & Privacy Advice? Measuring the Ability of LLMs to Refute Misconceptions
Chen, Y.; Arunasalam, A.; and Celik, Z. B. 2023 · 2023
Cited alongside, same era.
Closest in time.
Jailbreaking LLMs with Arabic Transliteration and Arabizi
Ghanim, M. A.; Almohaimeed, S.; Zheng, M.; Solihin, Y.; and Lou, Q. 2024 · 2024
Closest in time.
AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
Ghosh, S.; Varshney, P.; Galinkin, E.; and Parisien, C. 2024 · 2024
Closest in time.
Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment
Guo, Y.; Cui, G.; Yuan, L.; Ding, N.; Sun, Z.; Sun, B.; Chen, H.; Xie, R.; Zhou, J.; Lin, Y.; Liu, Z.; and Sun, M. 2024 · 2024
Closest in time.
WalledEval: A Comprehensive Safety Evaluation Toolkit for Large Language Models
Gupta, P.; Yau, L. Q.; Low, H. H.; Lee, I.-S.; Lim, H. M.; Teoh, Y. X.; Hng, K. J.; Liew, D. W.; Bhardwaj, R.; Bhardwaj, R.; and Poria, S. 2024a · 2024
Closest in time.
Evaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ
Holtermann, C.; Röttger, P.; Dill, T.; and Lauscher, A. 2024 · 2024
Closest in time.
Flames: Benchmarking Value Alignment of LLMs in Chinese
Huang, K.; Liu, X.; Guo, Q.; Sun, T.; Sun, J.; Wang, Y.; Zhou, Z.; Wang, Y.; Teng, Y.; Qiu, X.; Wang, Y.; and Lin, D. 2024a · 2024
Closest in time.
CBBQ: A Chinese Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models
Huang, Y.; and Xiong, D. 2024 · 2024
Closest in time.
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
Jiang, L.; Rao, K.; Han, S.; Ettinger, A.; Brahman, F.; Kumar, S.; Mireshghallah, N.; Lu, X.; Sap, M.; Choi, Y.; and Dziri, N. 2024 · 2024
Closest in time.
Multilingual Trolley Problems for Language Models
Jin, Z.; Kleiman-Weiner, M.; Piatti, G.; Levine, S.; Liu, J.; Adauto, F. G.; Ortu, F.; Strausz, A.; Sachan, M.; Mihalcea, R.; Choi, Y.; and Schölkopf, B. 2024b · 2024
Closest in time.
HELM Safety: Towards Standardized Safety Evaluations of Language Models
Kaiyom, F.; Ahmed, A.; Mai, Y.; Klyman, K.; Bommasani, R.; and Liang, P. 2024 · 2024
Closest in time.
The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
Kirk, H. R.; Whitefield, A.; Röttger, P.; Bean, A. M.; Margatina, K.; Mosquera, R.; Ciro, J. M.; Bartolo, M.; Williams, A.; He, H.; Vidgen, B.; and Hale, S. A. 2024 · 2024
Closest in time.
RewardBench: Evaluating Reward Models for Language Modeling
Lambert, N.; Pyatkin, V.; Morrison, J.; Miranda, L.; Lin, B. Y.; Chandu, K.; Dziri, N.; Kumar, S.; Zick, T.; Choi, Y.; Smith, N. A.; and Hajishirzi, H. 2024 · 2024
Closest in time.
KorNAT: LLM Alignment Benchmark for Korean Social Values and Common Knowledge
Lee, J.; Kim, M.; Kim, S.; Kim, J.; Won, S.; Lee, H.; and Choi, E. 2024 · 2024
Closest in time.
Gender Bias in Decision-Making with Large Language Models: A Study of Relationship Conflicts
Levy, S.; Adler, W.; Karver, T. S.; Dredze, M.; and Kaufman, M. R. 2024 · 2024
Closest in time.
SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
Li, L.; Dong, B.; Wang, R.; Hu, X.; Zuo, W.; Lin, D.; Qiao, Y.; and Shao, J. 2024a · 2024
Closest in time.
Stable bias: Evaluating societal representations in diffusion models
Luccioni, S.; Akiki, C.; Mitchell, M.; and Jernite, Y. 2024 · 2024
Closest in time.
Social Bias Probing: Fairness Benchmarking for Language Models
Manerba, M.; Stanczak, K.; Guidotti, R.; and Augenstein, I. 2024 · 2024
Closest in time.
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Mazeika, M.; Phan, L.; Yin, X.; Zou, A.; Wang, Z.; Mu, N.; Sakhaee, E.; Li, N.; Basart, S.; Li, B.; Forsyth, D.; and Hendrycks, D. 2024 · 2024
Closest in time.
The Llama 3 Herd of Models
Meta. 2024 · 2024
Closest in time.
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
Mireshghallah, N.; Kim, H.; Zhou, X.; Tsvetkov, Y.; Sap, M.; Shokri, R.; and Choi, Y. 2024 · 2024
Closest in time.
SG-Bench: Evaluating LLM Safety Generalization Across Diverse Tasks and Prompt Types
Mou, Y.; Zhang, S.; and Ye, W. 2024 · 2024
Closest in time.
Mu, N.; Chen, S.; Wang, Z.; Chen, S.; Karamardian, D.; Aljeraisy, L.; Alomair, B.; Hendrycks, D.; and Wagner, D. 2024 · 2024
Closest in time.
Women Are Beautiful, Men Are Leaders: Gender Stereotypes in Machine Translation and Language Modeling
Pikuliak, M.; Oresko, S.; Hrckova, A.; and Simko, M. 2024 · 2024
Closest in time.
CIVICS: Building a Dataset for Examining Culturally-Informed Values in Large Language Models
Pistilli, G.; Leidinger, A.; Jernite, Y.; Kasirzadeh, A.; Luccioni, A. S.; and Mitchell, M. 2024 · 2024
Closest in time.
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Qi, X.; Zeng, Y.; Xie, T.; Chen, P.-Y.; Jia, R.; Mittal, P.; and Henderson, P. 2024 · 2024
Closest in time.
BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices
Reuel, A.; Hardy, A.; Smith, C.; Lamparth, M.; Hardy, M.; and Kochenderfer, M. 2024 · 2024
Closest in time.
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Röttger, P.; Kirk, H.; Vidgen, B.; Attanasio, G.; Bianchi, F.; and Hovy, D. 2024 · 2024
Closest in time.
Towards Understanding Sycophancy in Language Models
Sharma, M.; Tong, M.; Korbak, T.; Duvenaud, D.; Askell, A.; Bowman, S. R.; DURMUS, E.; Hatfield-Dodds, Z.; Johnston, S. R.; Kravec, S. M.; Maxwell, T.; McCandlish, S.; Ndousse, K.; Rausch, O.; Schiefer, N.; Yan, D.; Zhang, M.; and Perez, E. 2024 · 2024
Closest in time.
” do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models
Shen, X.; Chen, Z.; Backes, M.; Shen, Y.; and Zhang, Y. 2024 · 2024
Closest in time.
Navigating the overkill in large language models
Shi, C.; Wang, X.; Ge, Q.; Gao, S.; Yang, X.; Gui, T.; Zhang, Q.; Huang, X.; Zhao, X.; and Lin, D. 2024 · 2024
Closest in time.
A StrongREJECT for Empty Jailbreaks
Souly, A.; Lu, Q.; Bowen, D.; Trinh, T.; Hsieh, E.; Pandey, S.; Abbeel, P.; Svegliato, J.; Emmons, S.; Watkins, O.; and Toyer, S. 2024 · 2024
Closest in time.
Trustllm: Trustworthiness in large language models
Sun, L.; Huang, Y.; Wang, H.; Wu, S.; Zhang, Q.; Gao, C.; Huang, Y.; Lyu, W.; Zhang, Y.; Li, X.; et al. 2024 · 2024
Closest in time.
Towards massive multilingual holistic bias
Tan, X. E.; Hansanti, P.; Wood, C.; Yu, B.; Ropers, C.; and Costa-jussà, M. R. 2024 · 2024
Closest in time.
ALERT: A Comprehensive Benchmark for Assessing Large Language Models’ Safety through Red Teaming
Tedeschi, S.; Friedrich, F.; Schramowski, P.; Kersting, K.; Navigli, R.; Nguyen, H.; and Li, B. 2024 · 2024
Closest in time.
All Languages Matter: On the Multilingual Safety of LLMs
Wang, W.; Tu, Z.; Chen, C.; Yuan, Y.; Huang, J.-t.; Jiao, W.; and Lyu, M. 2024b · 2024
Closest in time.
Do-Not-Answer: Evaluating Safeguards in LLMs
Wang, Y.; Li, H.; Han, X.; Nakov, P.; and Baldwin, T. 2024c · 2024
Closest in time.
Sorry-bench: Systematically evaluating large language model safety refusal behaviors
Xie, T.; Qi, X.; Zeng, Y.; Huang, Y.; Sehwag, U. M.; Huang, K.; He, L.; Wei, B.; Li, D.; Sheng, Y.; et al. 2024 · 2024
Closest in time.
Yang, A.; Yang, B.; Hui, B.; Zheng, B.; Yu, B.; Zhou, C.; Li, C.; Li, C.; Liu, D.; Huang, F.; Dong, G.; Wei, H.; Lin, H.; Tang, J.; Wang, J.; Yang, J.; Tu, J.; Zhang, J.; Ma, J.; Yang, J.; Xu, J.; Zhou, J.; Bai, J.; He, J.; Lin, J.; Dang, K.; Lu, K.; Chen, K.; Yang, K.; Li, M.; Xue, M.; Ni, N.; Zhang, P.; Wang, P.; Peng, R.; Men, R.; Gao, R.; Lin, R.; Wang, S.; Bai, S.; Tan, S.; Zhu, T.; Li, T.; Liu, T.; Ge, W.; Deng, X.; Zhou, X.; Ren, X.; Zhang, X.; Wei, X.; Ren, X.; Liu, X.; Fan, Y.; Yao, Y.; Zhang, Y.; Wan, Y.; Chu, Y.; Liu, Y.; Cui, Z.; Zhang, Z.; Guo, Z.; and Fan, Z. 2024 · 2024
Closest in time.
CoSafe: Evaluating Large Language Model Safety in Multi-Turn Dialogue Coreference
Yu, E.; Li, J.; Liao, M.; Wang, S.; Zuchen, G.; Mi, F.; and Hong, L. 2024a · 2024
Closest in time.
CMoralEval: A Moral Evaluation Benchmark for Chinese Large Language Models
Yu, L.; Leng, Y.; Huang, Y.; Wu, S.; Liu, H.; Ji, X.; Zhao, J.; Song, J.; Cui, T.; Cheng, X.; Liutao, L.; and Xiong, D. 2024c · 2024
Closest in time.
Yuan, X.; Li, J.; Wang, D.; Chen, Y.; Mao, X.; Huang, L.; Xue, H.; Wang, W.; Ren, K.; and Wang, J. 2024 · 2024
Closest in time.
WorldValuesBench: A Large-Scale Benchmark Dataset for Multi-Cultural Value Awareness of Language Models
Zhao, W.; Mondal, D.; Tandon, N.; Dillion, D.; Gray, K.; and Gu, Y. 2024a · 2024
Closest in time.
LMSYS-1M: A Large-Scale Real-World LLM Conversation Dataset
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Li, T.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Li, Z.; Lin, Z.; Xing, E.; Gonzalez, J. E.; Stoica, I.; and Zhang, H. 2024 · 2024
Closest in time.
Are Large Pre-Trained Language Models Leaking Your Personal Information?
Huang, J.; Shao, H.; and Chang, K. C.-C. 2022 · 2047
Closest in time.
BBQ: A hand-built bias benchmark for question answering
Parrish, A.; Chen, A.; Nangia, N.; Padmakumar, V.; Phang, J.; Thompson, J.; Htut, P. M.; and Bowman, S. 2022 · 2086
Closest in time.