Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are susceptible to a variety of risks, from non-faithful output to biased and toxic generations.
Nuanced Metrics for Measuring Unintended Bias with Real Data for Text Classification
Borkan, D.; Dixon, L.; Sorensen, J.; Thain, N.; and Vasserman, L. 2019 · 1903
Earlier work this paper cites.
Semantic noise matters for neural natural language generation
Dušek, O.; Howcroft, D. M.; and Rieser, V. 2019 · 1911
Earlier work this paper cites.
Sorting things out
Bowker, G. C.; and Star, S. L. 1999 · 1999
Earlier work this paper cites.
Twenty Newsgroups
Mitchell, T. 1999 · 1999
Earlier work this paper cites.
Machine-learning applications of algorithmic randomness
Vovk, V.; Gammerman, A.; and Saunders, C. 1999 · 1999
Earlier work this paper cites.
StereoSet: Measuring stereotypical bias in pretrained language models
Nadeem, M.; Bethke, A.; and Reddy, S. 2020 · 2004
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Gehman, S.; Gururangan, S.; Sap, M.; Choi, Y.; and Smith, N. A. 2020 · 2009
Earlier work this paper cites.
Stigma Management Communication: A Theory and Agenda for Applied Research on How Individuals Manage Moments of Stigmatized Identity
Meisenbach, R. J. 2010 · 2010
Earlier work this paper cites.
Measuring accessibility: positive and normative implementations of various accessibility indicators
Páez, A.; Scott, D. M.; and Morency, C. 2012 · 2012
Earlier work this paper cites.
INTRODUCTION , 3–6
CORRIGAN, P. W. 2014 · 2014
Earlier work this paper cites.
UNDERSTANDING STIGMA , 9–34
JONES, N.; and CORRIGAN, P. W. 2014 · 2014
Earlier work this paper cites.
Truth Is a Lie: Crowd Truth and the Seven Myths of Human Annotation
Aroyo, L.; and Welty, C. 2015 · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Bowman, S. R.; Angeli, G.; Potts, C.; and Manning, C. D. 2015 · 2015
Earlier work this paper cites.
#thyghgapp: Instagram Content Moderation and Lexical Variation in Pro-Eating Disorder Communities
Chancellor, S.; Pater, J. A.; Clear, T.; Gilbert, E.; and De Choudhury, M. 2016 · 2016
Earlier work this paper cites.
”Why should I trust you?” Explaining the predictions of any classifier
Ribeiro, M. T.; Singh, S.; and Guestrin, C. 2016 · 2016
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B.; Pritzel, A.; and Blundell, C. 2017 · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Lundberg, S. M.; and Lee, S.-I. 2017 · 2017
Earlier work this paper cites.
Axiomatic Attribution for Deep Networks
Sundararajan, M.; Taly, A.; and Yan, Q. 2017 · 2017
Earlier work this paper cites.
NewsQA: A Machine Comprehension Dataset
Trischler, A.; Wang, T.; Yuan, X.; Harris, J.; Sordoni, A.; Bachman, P.; and Suleman, K. 2017 · 2017
Earlier work this paper cites.
Data Statements for Natural Language Processing: Toward Mitigating System Bias and Enabling Better Science
Bender, E. M.; and Friedman, B. 2018 · 2018
Earlier work this paper cites.
Content or context moderation?
Caplan, R. 2018 · 2018
Earlier work this paper cites.
Situating methods in the magic of Big Data and AI
Elish, M. C.; and danah boyd. 2018 · 2018
Earlier work this paper cites.
Custodians of the Internet: Platforms, Content Moderation, and the Hidden Decisions That Shape Social Media
Gillespie, T. 2018 · 2018
Earlier work this paper cites.
The Burden of Stigma on Health and Well-Being: A Taxonomy of Concealment, Course, Disruptiveness, Aesthetics, Origin, and Peril Across 93 Stigmas
Pachankis, J. E.; Hatzenbuehler, M. L.; Wang, K.; Burton, C. L.; Crawford, F. W.; Phelan, J. C.; and Link, B. G. 2018 · 2018
Earlier work this paper cites.
Trust in Data Science: Collaboration, Translation, and Accountability in Corporate Data Science Projects
Passi, S.; and Jackson, S. J. 2018 · 2018
Earlier work this paper cites.
Detecting Egregious Conversations between Customers and Virtual Agents
Sandbank, T.; Shmueli-Scheuer, M.; Herzig, J.; Konopnicki, D.; Richards, J.; and Piorkowski, D. 2018 · 2018
Earlier work this paper cites.
Mind the GAP: A Balanced Corpus of Gendered Ambiguous Pronouns
Webster, K.; Recasens, M.; Axelrod, V.; and Baldridge, J. 2018 · 2018
Earlier work this paper cites.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Williams, A.; Nangia, N.; and Bowman, S. 2018 · 2018
Earlier work this paper cites.
HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Yang, Z.; Qi, P.; Zhang, S.; Bengio, Y.; Cohen, W.; Salakhutdinov, R.; and Manning, C. D. 2018 · 2018
Earlier work this paper cites.
Dominant Cultural and Personal Stigma Beliefs and the Utilization of Mental Health Services: A Cross-National Comparison
Bracke, P.; Delaruelle, K.; and Verhaeghe, M. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Earlier work this paper cites.
Ghost work: How to stop Silicon Valley from building a new global underclass
Gray, M. L.; and Suri, S. 2019 · 2019
Earlier work this paper cites.
The importance of context and intent in content moderation
Leetaru, K. 2019 · 2019
Earlier work this paper cites.
A simple recipe towards reducing hallucination in neural surface realisation
Nie, F.; Yao, J.-G.; Wang, J.; Pan, R.; and Lin, C.-Y. 2019 · 2019
Earlier work this paper cites.
BigBlueBot: teaching strategies for successful human-agent interactions
Weisz, J. D.; Jain, M.; Joshi, N. N.; Johnson, J.; and Lange, I. 2019 · 2019
Earlier work this paper cites.
Generating Hierarchical Explanations on Text Classification via Feature Interaction Detection
Chen, H.; Zheng, G.; and Ji, Y. 2020 · 2020
Earlier work this paper cites.
Data Feminism
D’Ignazio, C.; and Klein, L. F. 2020 · 2020
Earlier work this paper cites.
Social Chemistry 101: Learning to Reason about Social and Moral Norms
Forbes, M.; Hwang, J. D.; Shwartz, V.; Sap, M.; and Choi, Y. 2020 · 2020
Earlier work this paper cites.
Interpretation of NLP models through input marginalization
Kim, S.; Yi, J.; Kim, E.; and Yoon, S. 2020 · 2020
Earlier work this paper cites.
Conducting HCI Research on Stigma
Maestre, J. F. 2020 · 2020
Earlier work this paper cites.
Between Subjectivity and Imposition: Power Dynamics in Data Annotation for Computer Vision
Miceli, M.; Schuessler, M.; and Yang, T. 2020 · 2020
Earlier work this paper cites.
CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models
Nangia, N.; Vania, C.; Bhalerao, R.; and Bowman, S. R. 2020 · 2020
Earlier work this paper cites.
Classification with valid and adaptive coverage
Romano, Y.; Sesia, M.; and Candes, E. 2020 · 2020
Earlier work this paper cites.
How do Data Science Workers Collaborate? Roles, Workflows, and Tools
Zhang, A. X.; Muller, M.; and Wang, D. 2020 · 2020
Cited alongside, same era.
Machine Learning Model Drift Detection Via Weak Data Slices
Ackerman, S.; Dube, P.; Farchi, E.; Raz, O.; and Zalmanovici, M. 2021 · 2021
Cited alongside, same era.
Uncertainty Sets for Image Classifiers using Conformal Prediction
Angelopoulos, A. N.; Bates, S.; Jordan, M.; and Malik, J. 2021 · 2021
Cited alongside, same era.
Ground-Truth, Whose Truth? – Examining the Challenges with Annotating Toxic Text Datasets
Arhin, K.; Baldini, I.; Wei, D.; Ramamurthy, K. N.; and Singh, M. 2021 · 2021
Cited alongside, same era.
BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation
Dhamala, J.; Sun, T.; Kumar, V.; Krishna, S.; Pruksachatkun, Y.; Chang, K.; and Gupta, R. 2021 · 2021
Cited alongside, same era.
Judging facts, judging norms: Training machine learning models to judge humans requires a modified approach to labeling data
Balagoplan, A.; Madras, D.; Yang, D.; Hadfield-Menell, D.; Hadfield, G.; and Ghassemi, M. 2023 · 2023
Later among the works it cites.
Building AI for business: IBM’s Granite foundation models. 2023
2023
Later among the works it cites.
Can Large Language Models Be an Alternative to Human Evaluations?
Chiang, C.-H.; and Lee, H.-y. 2023 · 2023
Later among the works it cites.
A Survey on In-context Learning
Dong, Q.; Li, L.; Dai, D.; Zheng, C.; Wu, Z.; Chang, B.; Sun, X.; Xu, J.; Li, L.; and Sui, Z. 2023 · 2023
Later among the works it cites.
Influence Based Approaches to Algorithmic Fairness: A Closer Look
Ghosh, S.; Sattigeri, P.; Padhi, I.; Nagireddy, M.; and Chen, J. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Latent Hatred: A Benchmark for Understanding Implicit Hate Speech
ElSherief, M.; Ziems, C.; Muchlinski, D.; Anupindi, V.; Seybolt, J.; De Choudhury, M.; and Yang, D. 2021a · 2021
Cited alongside, same era.
Latent Hatred: A Benchmark for Understanding Implicit Hate Speech
ElSherief, M.; Ziems, C.; Muchlinski, D.; Anupindi, V.; Seybolt, J.; De Choudhury, M.; and Yang, D. 2021b · 2021
Cited alongside, same era.
Generate your counterfactuals: Towards controlled counterfactual generation for text
Madaan, N.; Padhi, I.; Panwar, N.; and Saha, D. 2021 · 2021
Cited alongside, same era.
Documenting Computer Vision Datasets: An Invitation to Reflexive Data Practices
Miceli, M.; Yang, T.; Naudts, L.; Schuessler, M.; Serbanescu, D.; and Hanna, A. 2021 · 2021
Cited alongside, same era.
Designing Ground Truth and the Social Life of Labels
Muller, M.; Wolf, C. T.; Andres, J.; Desmond, M.; Joshi, N. N.; Ashktorab, Z.; Sharma, A.; Brimijoin, K.; Pan, Q.; Duesterwald, E.; and Dugan, C. 2021 · 2021
Cited alongside, same era.
Mitigating harm in language models with conditional-likelihood filtration
Ngo, H.; Raterink, C.; Araújo, J. G.; Zhang, I.; Chen, C.; Morisot, A.; and Frosst, N. 2021 · 2021
Cited alongside, same era.
Uncertainty quantification and deep ensembles
Rahaman, R.; et al. 2021 · 2021
Cited alongside, same era.
RADAR: Robust AI-Text Detection via Adversarial Learning
Hu, X.; Chen, P.-Y.; and Ho, T.-Y. 2023 · 2023
Later among the works it cites.
IBM AI Risk Atlas. 2023
2023
Later among the works it cites.
IBM watsonx - An AI and data platform built for business. 2023
2023
Later among the works it cites.
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Inan, H.; Upasani, K.; Chi, J.; Rungta, R.; Iyer, K.; Mao, Y.; Tontchev, M.; Hu, Q.; Fuller, B.; Testuggine, D.; and Khabsa, M. 2023 · 2023
Later among the works it cites.
Baseline Defenses for Adversarial Attacks Against Aligned Language Models
Jain, N.; Schwarzschild, A.; Wen, Y.; Somepalli, G.; Kirchenbauer, J.; Chiang, P.; Goldblum, M.; Saha, A.; Geiping, J.; and Goldstein, T. 2023 · 2023
Later among the works it cites.
Jiang, A. Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D. S.; Casas, D. d. l.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; et al. 2023 · 2023
Later among the works it cites.
Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
Kim, S.; Shin, J.; Cho, Y.; Jang, J.; Longpre, S.; Lee, H.; Yun, S.; Shin, S.; Kim, S.; Thorne, J.; and Seo, M. 2023 · 2023
Later among the works it cites.
Unveiling Safety Vulnerabilities of Large Language Models
Kour, G.; Zalmanovici, M.; Zwerdling, N.; Goldbraich, E.; Fandina, O. N.; Anaby-Tavor, A.; Raz, O.; and Farchi, E. 2023 · 2023
Later among the works it cites.
Certifying llm safety against adversarial prompting
Kumar, A.; Agarwal, C.; Srinivas, S.; Feizi, S.; and Lakkaraju, H. 2023 · 2023
Later among the works it cites.
Black and Latinx Primary Caregiver Considerations for Developing and Implementing a Machine Learning–Based Model for Detecting Child Abuse and Neglect With Implications for Racial Bias Reduction: Qualitative Interview Study With Primary Caregivers
Landau, A. Y.; Blanchard, A.; Atkins, N.; Salazar, S.; Cato, K.; Patton, D. U.; and Topaz, M. 2023 · 2023
Later among the works it cites.
MISMATCH: Fine-grained Evaluation of Machine-generated Text with Mismatch Error Types
Murugesan, K.; Swaminathan, S.; Dan, S.; CHAUDHURY, S.; Gunasekara, C.; Crouse, M.; Mahajan, D.; Abdelaziz, I.; Fokoue, A.; Kapanipathi, P.; et al. 2023 · 2023
Later among the works it cites.
Scalable Extraction of Training Data from (Production) Language Models
Nasr, M.; Carlini, N.; Hayase, J.; Jagielski, M.; Cooper, A. F.; Ippolito, D.; Choquette-Choo, C. A.; Wallace, E.; Tramèr, F.; and Lee, K. 2023 · 2023
Later among the works it cites.
Applying Reflexivity to Artificial Intelligence for Researching Marginalized Communities and Real-World Problems
Nathan, A.; Aviv, L.; Siva, M.; Alex, A.; Sarah, D.; and Desmond, P. 2023 · 2023
Later among the works it cites.
Quantitative AI Risk Assessments: Opportunities and Challenges
Piorkowski, D.; Hind, M.; and Richards, J. 2023 · 2023
Later among the works it cites.
Supporting Human-AI Collaboration in Auditing LLMs with LLMs
Rastogi, C.; Tulio Ribeiro, M.; King, N.; Nori, H.; and Amershi, S. 2023 · 2023
Later among the works it cites.
NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails
Rebedea, T.; Dinu, R.; Sreedhar, M. N.; Parisien, C.; and Cohen, J. 2023 · 2023
Later among the works it cites.
Smoothllm: Defending large language models against jailbreaking attacks
Robey, A.; Wong, E.; Hassani, H.; and Pappas, G. J. 2023 · 2023
Later among the works it cites.
From Words to Watts: Benchmarking the Energy Costs of Large Language Model Inference
Samsi, S.; Zhao, D.; McDonald, J.; Li, B.; Michaleas, A.; Jones, M.; Bergeron, W.; Kepner, J.; Tiwari, D.; and Gadepally, V. 2023 · 2023
Later among the works it cites.
The Tail Wagging the Dog: Dataset Construction Biases of Social Bias Benchmarks
Selvam, N.; Dev, S.; Khashabi, D.; Khot, T.; and Chang, K.-W. 2023 · 2023
Later among the works it cites.
Large Language Models are Not Yet Human-Level Evaluators for Abstractive Summarization
Shen, C.; Cheng, L.; Nguyen, X.-P.; You, Y.; and Bing, L. 2023a · 2023
Later among the works it cites.
Muted: Multilingual Targeted Offensive Speech Identification and Visualization
Tillmann, C.; Trivedi, A.; Rosenthal, S.; Borse, S.; Zhang, R.; Sil, A.; and Bhattacharjee, B. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023 · 2023
Later among the works it cites.
Neural Architecture Search for Effective Teacher-Student Knowledge Transfer in Language Models
Trivedi, A.; Udagawa, T.; Merler, M.; Panda, R.; El-Kurdi, Y.; and Bhattacharjee, B. 2023 · 2023
Later among the works it cites.
Large Language Models are not Fair Evaluators
Wang, P.; Li, L.; Chen, L.; Cai, Z.; Zhu, D.; Lin, B.; Cao, Y.; Liu, Q.; Liu, T.; and Sui, Z. 2023 · 2023
Later among the works it cites.
Jailbroken: How Does LLM Safety Training Fail?
Wei, A.; Haghtalab, N.; and Steinhardt, J. 2023 · 2023
Later among the works it cites.
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E. P.; Zhang, H.; Gonzalez, J. E.; and Stoica, I. 2023 · 2023
Later among the works it cites.
JudgeLM: Fine-tuned Large Language Models are Scalable Judges
Zhu, L.; Wang, X.; and Wang, X. 2023 · 2023
Later among the works it cites.
Universal and Transferable Adversarial Attacks on Aligned Language Models
Zou, A.; Wang, Z.; Kolter, J. Z.; and Fredrikson, M. 2023 · 2023
Later among the works it cites.
Building Guardrails for Large Language Models
Dong, Y.; Mu, R.; Jin, G.; Qi, Y.; Hu, J.; Zhao, X.; Meng, J.; Ruan, W.; and Huang, X. 2024 · 2024
Closest in time.
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
Hu, X.; Chen, P.-Y.; and Ho, T.-Y. 2024 · 2024
Closest in time.
Foundation models: Opportunities, risks and mitigations
IBM AI Ethics Board. 2024 · 2024
Closest in time.
Language Models in Dialogue: Conversational Maxims for Human-AI Interactions
Miehling, E.; Nagireddy, M.; Sattigeri, P.; Daly, E. M.; Piorkowski, D.; and Richards, J. T. 2024 · 2024
Closest in time.
SocialStigmaQA: A Benchmark to Uncover Stigma Amplification in Generative Language Models
Nagireddy, M.; Chiazor, L.; Singh, M.; and Baldini, I. 2024a · 2024
Closest in time.
Multi-Level Explanations for Generative Language Models
Paes, L. M.; Wei, D.; Do, H. J.; Strobelt, H.; Luss, R.; Dhurandhar, A.; Nagireddy, M.; Ramamurthy, K. N.; Sattigeri, P.; Geyer, W.; and Ghosh, S. 2024 · 2024
Closest in time.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Qi, X.; Zeng, Y.; Xie, T.; Chen, P.-Y.; Jia, R.; Mittal, P.; and Henderson, P. 2024 · 2024
Closest in time.
TrustLLM: Trustworthiness in Large Language Models
Sun, L.; Huang, Y.; Wang, H.; Wu, S.; Zhang, Q.; Gao, C.; Huang, Y.; Lyu, W.; Zhang, Y.; Li, X.; Liu, Z.; Liu, Y.; Wang, Y.; Zhang, Z.; Kailkhura, B.; Xiong, C.; Zhang, C.; Xiao, C.; Li, C.; Xing, E.; Huang, F.; Liu, H.; Ji, H.; Wang, H.; Zhang, H.; Yao, H.; Kellis, M.; Zitnik, M.; Jiang, M.; Bansal, M.; Zou, J.; Pei, J.; Liu, J.; Gao, J.; Han, J.; Zhao, J.; Tang, J.; Wang, J.; Mitchell, J.; Shu, K.; Xu, K.; Chang, K.-W.; He, L.; Huang, L.; Backes, M.; Gong, N. Z.; Yu, P. S.; Chen, P.-Y.; Gu, Q.; Xu, R.; Ying, R.; Ji, S.; Jana, S.; Chen, T.; Liu, T.; Zhou, T.; Wang, W.; Li, X.; Zhang, X.; Wang, X.; Xie, X.; Chen, X.; Wang, X.; Liu, Y.; Ye, Y.; Cao, Y.; and Zhao, Y. 2024 · 2024
Closest in time.
Zeng, Y.; Lin, H.; Zhang, J.; Yang, D.; Jia, R.; and Shi, W. 2024 · 2024
Closest in time.