Fetching the paper…
Reading the bibliography…
Ensuring the trustworthiness of large language models (LLMs) is crucial.
Frequency principle: Fourier analysis sheds light on deep neural networks
Zhi-Qin John Xu, Yaoyu Zhang, Tao Luo, Yanyang Xiao, and Zheng Ma. 2019 · 1901
Earlier work this paper cites.
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R Bowman, Dipanjan Das, et al. 2019 · 1905
Earlier work this paper cites.
Elements of information theory
Thomas M Cover. 1999 · 1999
Earlier work this paper cites.
Estimating mutual information
Alexander Kraskov, Harald Stögbauer, and Peter Grassberger. 2004 · 2004
Earlier work this paper cites.
Measuring statistical dependence with hilbert-schmidt norms
Arthur Gretton, Olivier Bousquet, Alex Smola, and Bernhard Schölkopf. 2005 · 2005
Earlier work this paper cites.
A taxonomy of privacy
Daniel J Solove. 2005 · 2005
Earlier work this paper cites.
Trade-offs between membership privacy & adversarially robust learning
Jamie Hayes. 2020 · 2006
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky. 2015 · 2015
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
Deep variational information bottleneck
Alexander A Alemi, Ian Fischer, Joshua V Dillon, and Kevin Murphy. 2016 · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. 2016 · 2016
Earlier work this paper cites.
A closer look at memorization in deep networks
Devansh Arpit, Stanisław Jastrzębski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al. 2017 · 2017
Earlier work this paper cites.
Opening the black box of deep neural networks via information
Ravid Shwartz-Ziv and Naftali Tishby. 2017 · 2017
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018 · 2018
Earlier work this paper cites.
Ethics guidelines for trustworthy AI
European Commission, Content Directorate-General for Communications Networks, and Technology. 2019 · 2019
Earlier work this paper cites.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D Manning. 2019 · 2019
Earlier work this paper cites.
Do deep neural networks learn shallow learnable examples first?
Karttikeya Mangalam and Vinay Uday Prabhu. 2019 · 2019
Earlier work this paper cites.
Scalable mutual information estimation using dependence graphs
Morteza Noshad, Yu Zeng, and Alfred O Hero. 2019 · 2019
Earlier work this paper cites.
On variational bounds of mutual information
Ben Poole, Sherjil Ozair, Aaron Van Den Oord, Alex Alemi, and George Tucker. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
On the information bottleneck theory of deep learning
Andrew M Saxe, Yamini Bansal, Joel Dapello, Madhu Advani, Artemy Kolchinsky, Brendan D Tracey, and David D Cox. 2019 · 2019
Earlier work this paper cites.
How does bert answer questions? a layer-wise analysis of transformer representations
Betty Van Aken, Benjamin Winter, Alexander Löser, and Felix A Gers. 2019 · 2019
Earlier work this paper cites.
The information bottleneck problem and its applications in machine learning
Ziv Goldfeld and Yury Polyanskiy. 2020 · 2020
Earlier work this paper cites.
Investigating learning dynamics of BERT fine-tuning
Yaru Hao, Li Dong, Furu Wei, and Ke Xu. 2020 · 2020
Earlier work this paper cites.
Ro{bert}a: A robustly optimized {bert} pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Earlier work this paper cites.
The hsic bottleneck: Deep learning without back-propagation
Wan-Duo Kurt Ma, JP Lewis, and W Bastiaan Kleijn. 2020 · 2020
Earlier work this paper cites.
What happens to bert embeddings during fine-tuning?
Amil Merchant, Elahe Rahimtoroghi, Ellie Pavlick, and Ian Tenney. 2020 · 2020
Earlier work this paper cites.
On the interplay between fine-tuning and sentence-level probing for linguistic knowledge in pre-trained transformers
Marius Mosbach, Anna Khokhlova, Michael A Hedderich, and Dietrich Klakow. 2020 · 2020
Earlier work this paper cites.
Information-theoretic probing for linguistic structure
Tiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod, Adina Williams, and Ryan Cotterell. 2020 · 2020
Earlier work this paper cites.
Intrinsic dimensionality explains the effectiveness of language model fine-tuning
Armen Aghajanyan, Sonal Gupta, and Luke Zettlemoyer. 2021 · 2021
Cited alongside, same era.
Proposal for a regulation of the european parliament and of the council laying down harmonised rules on artificial intelligence (artificial intelligence act) and amending certain union legislative acts, pub. l. no. com(2021) 206 final
European Commission. 2021b · 2021
Cited alongside, same era.
On information plane analyses of neural network classifiers–a review
Bernhard C Geiger. 2021 · 2021
Cited alongside, same era.
Implicit representations of meaning in neural language models
Belinda Z Li, Maxwell Nye, and Jacob Andreas. 2021 · 2021
Cited alongside, same era.
On the impact of hard adversarial instances on overfitting in adversarial training
Pretraining language models with human preferences
Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Vinayak Bhalerao, Christopher Buckley, Jason Phang, Samuel R Bowman, and Ethan Perez. 2023 · 2023
Later among the works it cites.
Differential privacy has bounded impact on fairness in classification
Paul Mangold, Michaël Perrot, Aurélien Bellet, and Marc Tommasi. 2023 · 2023
Later among the works it cites.
Samuel Marks and Max Tegmark. 2023 · 2023
Later among the works it cites.
Niloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov, Maarten Sap, Reza Shokri, and Yejin Choi. 2023 · 2023
Later among the works it cites.
An emulator for fine-tuning large language models using small language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen Liu, Zhichao Huang, Mathieu Salzmann, Tong Zhang, and Sabine Süsstrunk. 2021 · 2021
Cited alongside, same era.
Information bottleneck: Exact analysis of (quantized) neural networks
Stephan Sloth Lorenzen, Christian Igel, and Mads Nielsen. 2021 · 2021
Cited alongside, same era.
StereoSet: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy. 2021 · 2021
Cited alongside, same era.
Probing the probing paradigm: Does probing accuracy entail task relevance?
Abhilasha Ravichander, Yonatan Belinkov, and Eduard Hovy. 2021 · 2021
Cited alongside, same era.
Adversarial glue: A multi-task benchmark for robustness evaluation of language models
Boxin Wang, Chejian Xu, Shuohang Wang, Zhe Gan, Yu Cheng, Jianfeng Gao, Ahmed Hassan Awadallah, and Bo Li. 2021 · 2021
Cited alongside, same era.
To be robust or to be fair: Towards fairness in adversarial training
Han Xu, Xiaorui Liu, Yaxin Li, Anil Jain, and Jiliang Tang. 2021 · 2021
Cited alongside, same era.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov. 2022 · 2022
Cited alongside, same era.
Discovering latent knowledge in language models without supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. 2022 · 2022
Cited alongside, same era.
Eric Mitchell, Rafael Rafailov, Archit Sharma, Chelsea Finn, and Christopher D. Manning. 2023 · 2023
Later among the works it cites.
Emergent linear representations in world models of self-supervised sequence models
Neel Nanda, Andrew Lee, and Martin Wattenberg. 2023 · 2023
Later among the works it cites.
A taxonomy of trustworthiness for artificial intelligence: Connecting properties of trustworthiness with risk management and the ai lifecycle
Jessica Newman. 2023 · 2023
Later among the works it cites.
The linear representation hypothesis and the geometry of large language models
Kiho Park, Yo Joong Choe, and Victor Veitch. 2023 · 2023
Later among the works it cites.
Toward transparent ai: A survey on interpreting the inner structures of deep neural networks
Tilman Räuker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell. 2023 · 2023
Later among the works it cites.
Steering llama 2 via contrastive activation addition
Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner. 2023 · 2023
Later among the works it cites.
Aviv Slobodkin, Omer Goldman, Avi Caciularu, Ido Dagan, and Shauli Ravfogel. 2023 · 2023
Later among the works it cites.
Principle-driven self-alignment of language models from scratch with minimal human supervision
Zhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang, Zhenfang Chen, David Cox, Yiming Yang, and Chuang Gan. 2023 · 2023
Later among the works it cites.
Artificial intelligence risk management framework (ai rmf 1.0)
Elham Tabassi. 2023 · 2023
Later among the works it cites.
Joma: Demystifying multilayer transformers via joint dynamics of mlp and attention
Yuandong Tian, Yiping Wang, Zhenyu Zhang, Beidi Chen, and Simon Du. 2023 · 2023
Later among the works it cites.
Activation addition: Steering language models without optimization
Alex Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, and Monte MacDiarmid. 2023 · 2023
Later among the works it cites.
Haoran Wang and Kai Shu. 2023 · 2023
Later among the works it cites.
Explainability for large language models: A survey
Haiyan Zhao, Hanjie Chen, Fan Yang, Ninghao Liu, Huiqi Deng, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, and Mengnan Du. 2023 · 2023
Later among the works it cites.
Promptbench: Towards evaluating the robustness of large language models on adversarial prompts
Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Neil Zhenqiang Gong, Yue Zhang, et al. 2023 · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al. 2023 · 2023
Later among the works it cites.
Taichi: Improving the robustness of nlp models by seeking common ground while reserving differences
Huimin Chen, Chengyu Wang, Yanhao Wang, Cen Chen, and Yinggui Wang. 2024 · 2024
Closest in time.
Olmo: Accelerating the science of language models
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al. 2024 · 2024
Closest in time.
Salad-bench: A hierarchical and comprehensive safety benchmark for large language models
Lijun Li, Bowen Dong, Ruohui Wang, Xuhao Hu, Wangmeng Zuo, Dahua Lin, Yu Qiao, and Jing Shao. 2024 · 2024
Closest in time.
Tuning language models by proxy
Alisa Liu, Xiaochuang Han, Yizhong Wang, Yulia Tsvetkov, Yejin Choi, and Noah A. Smith. 2024 · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024 · 2024
Closest in time.
Trustllm: Trustworthiness in large language models
Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, et al. 2024 · 2024
Closest in time.
Inferaligner: Inference-time alignment for harmlessness through cross-model guidance
Pengyu Wang, Dong Zhang, Linyang Li, Chenkun Tan, Xinghao Wang, Ke Ren, Botian Jiang, and Xipeng Qiu. 2024 · 2024
Closest in time.
Self-rewarding language models
Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Sainbayar Sukhbaatar, Jing Xu, and Jason Weston. 2024 · 2024
Closest in time.
Random smooth-based certified defense against text adversarial attack
Zeliang Zhang, Wei Yao, Susan Liang, and Chenliang Xu. 2024b · 2024
Closest in time.
Explaining generalization power of a dnn using interactive concepts
Huilin Zhou, Hao Zhang, Huiqi Deng, Dongrui Liu, Wen Shen, Shih-Han Chan, and Quanshi Zhang. 2024 · 2024
Closest in time.