Fetching the paper…
Reading the bibliography…
Ensuring awareness of fairness and privacy in Large Language Models (LLMs) is critical.
A definition of conditional mutual information for arbitrary ensembles
Aaron D Wyner. 1978 · 1978
Earlier work this paper cites.
Evaluation metrics for language models
Stanley F Chen, Douglas Beeferman, and Roni Rosenfeld. 1998 · 1998
Earlier work this paper cites.
Secure multi-party computation problems and their applications: a review and open problems
Wenliang Du and Mikhail J Atallah. 2001 · 2001
Earlier work this paper cites.
Mutual information theory for adaptive mixture models
Zheng Rong Yang and Mark Zwolinski. 2001 · 2001
Earlier work this paper cites.
Estimating mutual information
Alexander Kraskov, Harald Stögbauer, and Peter Grassberger. 2004 · 2004
Earlier work this paper cites.
Privacy in deep learning: A survey
Fatemehsadat Mireshghallah, Mohammadkazem Taram, Praneeth Vepakomma, Abhishek Singh, Ramesh Raskar, and Hadi Esmaeilzadeh. 2020 · 2004
Earlier work this paper cites.
Measuring statistical dependence with hilbert-schmidt norms
Arthur Gretton, Olivier Bousquet, Alex Smola, and Bernhard Schölkopf. 2005 · 2005
Earlier work this paper cites.
Differential privacy
Cynthia Dwork. 2006 · 2006
Earlier work this paper cites.
Calibrating noise to sensitivity in private data analysis
Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam Smith. 2006 · 2006
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Information theory
Robert B Ash. 2012 · 2012
Earlier work this paper cites.
Fairness through awareness
Cynthia Dwork, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Richard Zemel. 2012 · 2012
Earlier work this paper cites.
Striving for Simplicity: The All Convolutional Net
J Springenberg, Alexey Dosovitskiy, Thomas Brox, and M Riedmiller. 2015 · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks
Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg. 2016 · 2016
Earlier work this paper cites.
Counterfactual fairness
Matt J Kusner, Joshua Loftus, Chris Russell, and Ricardo Silva. 2017 · 2017
Earlier work this paper cites.
RACE: Large-scale ReAding comprehension dataset from examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017 · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. 2017 · 2017
Earlier work this paper cites.
Learning Important Features Through Propagating Activation Differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. 2017 · 2017
Earlier work this paper cites.
Axiomatic Attribution for Deep Networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018 · 2018
Earlier work this paper cites.
Fairness definitions explained
Sahil Verma and Julia Rubin. 2018 · 2018
Earlier work this paper cites.
Differential privacy has disparate impact on model accuracy
Eugene Bagdasaryan, Omid Poursaeed, and Vitaly Shmatikov. 2019 · 2019
Earlier work this paper cites.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
On the compatibility of privacy and fairness
Rachel Cummings, Varun Gupta, Dhamma Kimpara, and Jamie Morgenstern. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Designing and Interpreting Probes With Control Tasks
John Hewitt and Percy Liang. 2019 · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Earlier work this paper cites.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Earlier work this paper cites.
On variational bounds of mutual information
Ben Poole, Sherjil Ozair, Aaron Van Den Oord, Alex Alemi, and George Tucker. 2019 · 2019
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Earlier work this paper cites.
Privacy and security issues in deep learning: A survey
Ximeng Liu, Lehui Xie, Yaopeng Wang, Jian Zou, Jinbo Xiong, Zuobin Ying, and Athanasios V Vasilakos. 2020 · 2020
Earlier work this paper cites.
Differentially private representation for NLP: Formal guarantee and an empirical study on privacy and fairness
Lingjuan Lyu, Xuanli He, and Yitong Li. 2020 · 2020
Earlier work this paper cites.
The hsic bottleneck: Deep learning without back-propagation
Wan-Duo Kurt Ma, JP Lewis, and W Bastiaan Kleijn. 2020 · 2020
Earlier work this paper cites.
A Survey on Explainable Artificial Intelligence (XAI): Towards Medical XAI
Erico Tjoa and Cuntai Guan. 2020 · 2020
Cited alongside, same era.
Asking and answering questions to evaluate the factual consistency of summaries
Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020 · 2020
Cited alongside, same era.
Trade-offs between fairness and privacy in machine learning
Sushant Agarwal. 2021 · 2021
Cited alongside, same era.
Summeval: Re-evaluating summarization evaluation
Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Cited alongside, same era.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021 · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Later among the works it cites.
Function Vectors in Large Language Models
Eric Todd, Millicent L Li, Arnab Sen Sharma, Aaron Mueller, Byron C Wallace, and David Bau. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Later among the works it cites.
Label words are anchors: An information flow perspective for understanding in-context learning
Lean Wang, Lei Li, Damai Dai, Deli Chen, Hao Zhou, Fandong Meng, Jie Zhou, and Xu Sun. 2023a · 2023
Later among the works it cites.
Large language models illuminate a progressive pathway to artificial healthcare assistant: A review
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks
Rabeeh Karimi Mahabadi, Sebastian Ruder, Mostafa Dehghani, and James Henderson. 2021 · 2021
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Cited alongside, same era.
Compacter: Efficient low-rank hypercomplex adapter layers
Rabeeh Karimi mahabadi, James Henderson, and Sebastian Ruder. 2021 · 2021
Cited alongside, same era.
A survey on bias and fairness in machine learning
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021 · 2021
Cited alongside, same era.
Knowledge neurons in pretrained transformers
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2022 · 2022
Cited alongside, same era.
Toy models of superposition
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah. 2022 · 2022
Cited alongside, same era.
Mingze Yuan, Peng Bao, Jiajia Yuan, Yunhao Shen, Zifan Chen, Yi Xie, Jie Zhao, Yang Chen, Li Zhang, Lin Shen, et al. 2023 · 2023
Later among the works it cites.
Adaptive budget allocation for parameter-efficient fine-tuning
Qingru Zhang, Minshuo Chen, Alexander Bukharin, Pengcheng He, Yu Cheng, Weizhu Chen, and Tuo Zhao. 2023 · 2023
Later among the works it cites.
Chain-of-thought unfaithfulness as disguised accuracy
Oliver Bentham, Nathan Stringham, and Ana Marasovic. 2024 · 2024
Closest in time.
Fairness in machine learning: A survey
Simon Caton and Christian Haas. 2024 · 2024
Closest in time.
Or-bench: An over-refusal benchmark for large language models
Justin Cui, Wei-Lin Chiang, Ion Stoica, and Cho-Jui Hsieh. 2024 · 2024
Closest in time.
Explainable and interpretable multimodal large language models: A comprehensive survey
Yunkai Dang, Kaichen Huang, Jiahao Huo, Yibo Yan, Sirui Huang, Dongrui Liu, Mengxi Gao, Jie Zhang, Chen Qian, Kun Wang, et al. 2024 · 2024
Closest in time.
Covert malicious finetuning: Challenges in safeguarding LLM adaptation
Danny Halawi, Alexander Wei, Eric Wallace, Tony Tong Wang, Nika Haghtalab, and Jacob Steinhardt. 2024 · 2024
Closest in time.
DP-OPT: Make large language model your privacy-preserving prompt engineer
Junyuan Hong, Jiachen T. Wang, Chenhui Zhang, Zhangheng LI, Bo Li, and Zhangyang Wang. 2024 · 2024
Closest in time.
Vlsbench: Unveiling visual leakage in multimodal safety
Xuhao Hu, Dongrui Liu, Hao Li, Xuanjing Huang, and Jing Shao. 2024 · 2024
Closest in time.
A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Andrew Lee, Xiaoyan Bai, Itamar Pres, Martin Wattenberg, Jonathan K Kummerfeld, and Rada Mihalcea. 2024 · 2024
Closest in time.
Chaochao Lu, Chen Qian, Guodong Zheng, Hongxing Fan, Hongzhi Gao, Jie Zhang, Jing Shao, Jingyi Deng, Jinlan Fu, Kexin Huang, et al. 2024 · 2024
Closest in time.
From understanding to utilization: A survey on explainability for large language models
Haoyan Luo and Lucia Specia. 2024 · 2024
Closest in time.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. 2024 · 2024
Closest in time.
Towards tracing trustworthiness dynamics: Revisiting pre-training period of large language models
Chen Qian, Jie Zhang, Wei Yao, Dongrui Liu, Zhenfei Yin, Yu Qiao, Yong Liu, and Jing Shao. 2024 · 2024
Closest in time.
GPQA: A graduate-level google-proof q&a benchmark
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. 2024 · 2024
Closest in time.
Identifying semantic induction heads to understand in-context learning
Jie Ren, Qipeng Guo, Hang Yan, Dongrui Liu, Xipeng Qiu, and Dahua Lin. 2024 · 2024
Closest in time.
Steering llama 2 via contrastive activation addition
Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Turner. 2024 · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Adly Templeton. 2024 · 2024
Closest in time.
Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. 2024 · 2024
Closest in time.
The art of defending: A systematic evaluation and analysis of LLM defense strategies on safety and over-defensiveness
Neeraj Varshney, Pavel Dolin, Agastya Seth, and Chitta Baral. 2024 · 2024
Closest in time.
Assessing the brittleness of safety alignment via pruning and low-rank modifications
Boyi Wei, Kaixuan Huang, Yangsibo Huang, Tinghao Xie, Xiangyu Qi, Mengzhou Xia, Prateek Mittal, Mengdi Wang, and Peter Henderson. 2024 · 2024
Closest in time.
Reft: Representation finetuning for language models
Zhengxuan Wu, Aryaman Arora, Zheng Wang, Atticus Geiger, Dan Jurafsky, Christopher D Manning, and Christopher Potts. 2024 · 2024
Closest in time.
WizardLM: Empowering large pre-trained language models to follow complex instructions
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, Qingwei Lin, and Daxin Jiang. 2024 · 2024
Closest in time.
Reef: Representation encoding fingerprints for large language models
Jie Zhang, Dongrui Liu, Chen Qian, Linfeng Zhang, Yong Liu, Yu Qiao, and Jing Shao. 2024 · 2024
Closest in time.
Privacy-preserving large language models: Mechanisms, applications, and future directions
Guoshenghui Zhao and Eric Song. 2024 · 2024
Closest in time.
Llamafactory: Unified efficient fine-tuning of 100+ language models
Yaowei Zheng, Richong Zhang, Junhao Zhang, Yanhan Ye, Zheyan Luo, Zhangchi Feng, and Yongqiang Ma. 2024 · 2024
Closest in time.
Seer: Self-explainability enhancement of large language models’ representations
Guanxu Chen, Dongrui Liu, Tao Luo, and Jing Shao. 2025 · 2025
Closest in time.
Xiaoya Lu, Dongrui Liu, Yi Yu, Luxin Xu, and Jing Shao. 2025 · 2025
Closest in time.
Led-merging: Mitigating safety-utility conflicts in model merging with location-election-disjoint
Qianli Ma, Dongrui Liu, Qian Chen, Linfeng Zhang, and Jing Shao. 2025 · 2025
Closest in time.
Faitheval: Can your language model stay faithful to context, even if ”the moon is made of marshmallows”
Yifei Ming, Senthil Purushwalkam, Shrey Pandit, Zixuan Ke, Xuan-Phi Nguyen, Caiming Xiong, and Shafiq Joty. 2025 · 2025
Closest in time.
Towards a unified multi-dimensional evaluator for text generation
Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han. 2022 · 2038
Closest in time.