Fetching the paper…
Reading the bibliography…
We propose an auditing method to identify whether a large language model (LLM) encodes patterns such as hallucinations in its internal states, which may propagate to downstream tasks.
Goodness-of-fit test statistics that dominate the kolmogorov statistics
Robert H Berk and Douglas H Jones · 1979
Earlier work this paper cites.
Higher criticism for detecting sparse heterogeneous mixtures
David Donoho and Jiashun Jin · 2004
Earlier work this paper cites.
Mathematical statistics and data analysis
John A Rice · 2007
Earlier work this paper cites.
P-values are random variables
Duncan J Murdoch, Yu-Ling Tsai, and James Adcock · 2008
Earlier work this paper cites.
Fast subset scan for spatial pattern detection
Daniel B. Neill · 2012
Earlier work this paper cites.
Fast generalized subset scan for anomalous pattern detection
Edward McFowland, Skyler Speakman, and Daniel B Neill · 2013
Earlier work this paper cites.
Non-parametric scan statistics for event detection and forecasting in heterogeneous social media graphs
Feng Chen and Daniel B Neill · 2014
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai · 2016
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Ethical challenges in data-driven dialogue systems
Peter Henderson, Koustuv Sinha, Nicolas Angelard-Gontier, Nan Rosemary Ke, Genevieve Fried, Ryan Lowe, and Joelle Pineau · 2018
Earlier work this paper cites.
Edward McFowland III, Sriram Somanchi, and Daniel B Neill · 2018
Earlier work this paper cites.
Sentence encoders on stilts: Supplementary training on intermediate labeled-data tasks
Jason Phang, Thibault Févry, and Samuel R Bowman · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Evaluating the underlying gender bias in contextualized word embeddings
Christine Basta, Marta R Costa-jussà, and Noe Casas · 2019
Earlier work this paper cites.
Identifying and reducing gender bias in word-level language models
Shikha Bordia and Samuel Bowman · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova · 2019
Earlier work this paper cites.
Measuring bias in contextualized word representations
Keita Kurita, Nidhi Vyas, Ayush Pareek, Alan W Black, and Yulia Tsvetkov · 2019
Earlier work this paper cites.
On measuring social biases in sentence encoders
Chandler May, Alex Wang, Shikha Bordia, Samuel Bowman, and Rachel Rudinger · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
Assessing social and intersectional biases in contextualized word representations
Yi Chern Tan and L Elisa Celis · 2019
Cited alongside, same era.
Gender bias in contextualized word embeddings
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Ryan Cotterell, Vicente Ordonez, and Kai-Wei Chang · 2019
Cited alongside, same era.
Identifying audio adversarial examples via anomalous pattern detection
Victor Akinwande, Celia Cintas, Skyler Speakman, and Srihari Sridharan · 2020
Cited alongside, same era.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith · 2020
Cited alongside, same era.
Mixout: Effective regularization to finetune large-scale pretrained language models
Cheolhyoung Lee, Kyunghyun Cho, and Wanmo Kang · 2020
Cited alongside, same era.
What do toothbrushes do in the kitchen? how transformers think our world is structured
Alexander Henlein and Alexander Mehler · 2022
Later among the works it cites.
Out-of-distribution detection in dermatology using input perturbation and subset scanning
Hannah Kim, Girmaw Abebe Tadesse, Celia Cintas, Skyler Speakman, and Kush Varshney · 2022
Later among the works it cites.
The data-production dispositif
Milagros Miceli and Julian Posada · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models, 2022
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer · 2022
Later among the works it cites.
The internal state of an llm knows when its lying
Amos Azaria and Tom Mitchell · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards debiasing sentence representations
Paul Pu Liang, Irene Mengze Li, Emily Zheng, Yao Chong Lim, Ruslan Salakhutdinov, and Louis-Philippe Morency · 2020
Cited alongside, same era.
Crows-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R Bowman · 2020
Cited alongside, same era.
Towards debiasing nlu models from unknown biases
PA Utama, NS Moosavi, and I Gurevych · 2020
Cited alongside, same era.
Dialogpt: Large-scale generative pre-training for conversational response generation
Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and William B Dolan · 2020
Cited alongside, same era.
Better fine-tuning by reducing representational collapse
Armen Aghajanyan, Akshat Shrivastava, Anchit Gupta, Naman Goyal, Luke Zettlemoyer, and Sonal Gupta · 2021
Cited alongside, same era.
Extensive study on the underlying gender bias in contextualized word embeddings
Christine Basta, Marta R Costa-Jussa, and Noe Casas · 2021
Cited alongside, same era.
Detecting adversarial attacks via subset scanning of autoencoder activations and reconstruction error
Celia Cintas, Skyler Speakman, Victor Akinwande, William Ogallo, Komminist Weldemariam, Srihari Sridharan, and Edward McFowland · 2021
Cited alongside, same era.
Abeba Birhane, Vinay Prabhu, Sang Han, and Vishnu Naresh Boddeti · 2023
Closest in time.
Toxicity in chatgpt: Analyzing persona-assigned language models
Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan · 2023
Closest in time.
Building stereotype repositories with llms and community engagement for scale and depth
Sunipa Dev, Akshita Jha, Jaya Goyal, Dinesh Tewari, Shachi Dave, and Vinodkumar Prabhakaran · 2023
Closest in time.
A survey on large language models: Applications, challenges, limitations, and practical usage
Muhammad Usman Hadi, R Qureshi, A Shah, M Irfan, A Zafar, MB Shaikh, N Akhtar, J Wu, and S Mirjalili · 2023
Closest in time.
Chatgpt for shaping the future of dentistry: the potential of multi-modal large language model
Hanyao Huang, Ou Zheng, Dongdong Wang, Jiayi Yin, Zijin Wang, Shengxuan Ding, Heng Yin, Chuan Xu, Renjie Yang, Qian Zheng, et al · 2023
Closest in time.
Unveiling theory of mind in large language models: A parallel to single neurons in the human brain
Mohsen Jamali, Ziv M Williams, and Jing Cai · 2023
Closest in time.
Spatially constrained adversarial attack detection and localization in the representation space of optical flow networks
Hannah Kim, Celia Cintas, Girmaw Abebe Tadesse, and Skyler Speakman · 2023
Closest in time.
Still no lie detector for language models: Probing empirical and conceptual roadblocks
BA Levinstein and Daniel A Herrmann · 2023
Closest in time.
Trustworthy llms: a survey and guideline for evaluating large language models’ alignment
Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li · 2023
Closest in time.
Unmasking nationality bias: A study of human perception of nationalities in ai-generated articles
Pranav Narayanan Venkit, Sanjana Gautam, Ruchi Panchanadikar, Ting-Hao Huang, and Shomir Wilson · 2023
Closest in time.
Large language models in medicine
Arun James Thirunavukarasu, Darren Shu Jeng Ting, Kabilan Elangovan, Laura Gutierrez, Ting Fang Tan, and Daniel Shu Wei Ting · 2023
Closest in time.
Shoujie Tong, Heming Xia, Damai Dai, Tianyu Liu, Binghuai Lin, Yunbo Cao, and Zhifang Sui · 2023
Closest in time.
Harnessing the power of llms in practice: A survey on chatgpt and beyond
Jingfeng Yang, Hongye Jin, Ruixiang Tang, Xiaotian Han, Qizhang Feng, Haoming Jiang, Bing Yin, and Xia Hu · 2023
Closest in time.
Red teaming chatgpt via jailbreaking: Bias, robustness, reliability and toxicity
Terry Yue Zhuo, Yujin Huang, Chunyang Chen, and Zhenchang Xing · 2023
Closest in time.