Fetching the paper…
Reading the bibliography…
Explainable AI (XAI) refers to techniques that provide human-understandable insights into the workings of AI models.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
S3-Rec: Self-supervised learning for sequential recommendation with mutual information maximization. In CIKM . 1893–1902
Kun Zhou, Hui Wang, Wayne Xin Zhao, Yutao Zhu, Sirui Wang, Fuzheng Zhang, Zhongyuan Wang, and Ji-Rong Wen. 2020 · 1902
Earlier work this paper cites.
The principle of minimized iterations in the solution of the matrix eigenvalue problem
Walter Edwin Arnoldi. 1951 · 1951
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: A strategy employed by V1?
Bruno A Olshausen and David J Field. 1997 · 1997
Earlier work this paper cites.
Support vector machines
Marti A. Hearst, Susan T Dumais, Edgar Osuna, John Platt, and Bernhard Scholkopf. 1998 · 1998
Earlier work this paper cites.
The probabilistic relevance framework: BM25 and beyond
Stephen Robertson, Hugo Zaragoza, et al · 2009
Earlier work this paper cites.
Intelligible models for classification and regression. In SIGKDD . 150–158
Yin Lou, Rich Caruana, and Johannes Gehrke. 2012 · 2012
Earlier work this paper cites.
Distributional vectors encode referential attributes. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing . 12–21
Abhijeet Gupta, Gemma Boleda, Marco Baroni, and Sebastian Padó. 2015 · 2015
Earlier work this paper cites.
Decision tree methods: applications for classification and prediction
Yan-Yan Song and LU Ying. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition. In CVPR . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
beta-vae: Learning basic visual concepts with a constrained variational framework. In ICLR
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner. 2016 · 2016
Earlier work this paper cites.
Understanding neural networks through representation erasure
Jiwei Li, Will Monroe, and Dan Jurafsky. 2016b · 2016
Earlier work this paper cites.
" Why should i trust you?" Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining . 1135–1144
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Grad-CAM: Why did you say that?
Ramprasaath R Selvaraju, Abhishek Das, Ramakrishna Vedantam, Michael Cogswell, Devi Parikh, and Dhruv Batra. 2016 · 2016
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations. In CVPR . 6541–6549
David Bau, Bolei Zhou, Aditya Khosla, Aude Oliva, and Antonio Torralba. 2017 · 2017
Earlier work this paper cites.
What do neural machine translation models learn about morphology?
Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James Glass. 2017 · 2017
Earlier work this paper cites.
Unsupervised learning of disentangled representations from video
Emily L Denton et al · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim. 2017 · 2017
Earlier work this paper cites.
Translation-based recommendation. In RecSys . 161–169
Ruining He, Wang-Cheng Kang, and Julian McAuley. 2017 · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions. In ICML . PMLR, 1885–1894
Pang Wei Koh and Percy Liang. 2017 · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee. 2017 · 2017
Earlier work this paper cites.
Explaining nonlinear classification decisions with deep taylor decomposition
Grégoire Montavon, Sebastian Lapuschkin, Alexander Binder, Wojciech Samek, and Klaus-Robert Müller. 2017 · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences. In ICML . PMLR, 3145–3153
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. 2017 · 2017
Earlier work this paper cites.
Counterfactual explanations without opening the black box: Automated decisions and the GDPR
Sandra Wachter, Brent Mittelstadt, and Chris Russell. 2017 · 2017
Earlier work this paper cites.
Where is your evidence: Improving fact-checking by justification modeling. In Proceedings of the first workshop on fact extraction and verification (FEVER) . 85–90
Tariq Alhindi, Savvas Petridis, and Smaranda Muresan. 2018 · 2018
Earlier work this paper cites.
Linear algebraic structure of word senses, with applications to polysemy
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski. 2018 · 2018
Earlier work this paper cites.
GAN Dissection: Visualizing and Understanding Generative Adversarial Networks. In ICLR
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Bolei Zhou, Joshua B Tenenbaum, William T Freeman, and Antonio Torralba. 2018 · 2018
Earlier work this paper cites.
e-snli: Natural language inference with natural language explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018 · 2018
Earlier work this paper cites.
Explainable artificial intelligence: A survey. In 2018 41st International convention on information and communication technology, electronics and microelectronics (MIPRO) . IEEE, 0210–0215
Filip Karlo Došilović, Mario Brčić, and Nikica Hlupić. 2018 · 2018
Earlier work this paper cites.
HotFlip: White-Box Adversarial Examples for Text Classification. In Proceedings of the 56th ACL (Volume 2: Short Papers) . 31–36
Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2018 · 2018
Earlier work this paper cites.
Residual Connections Encourage Iterative Inference. In International Conference on Learning Representations
Stanisław Jastrzebski, Devansh Arpit, Nicolas Ballas, Vikas Verma, Tong Che, and Yoshua Bengio. 2018 · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav). In International conference on machine learning . PMLR, 2668–2677
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al · 2018
Earlier work this paper cites.
Vqa-e: Explaining, elaborating, and enhancing your answers for visual questions. In ECCV
Q. Li, Q. Tao, S. Joty, J. Cai, and J. Luo. 2018 · 2018
Earlier work this paper cites.
Adversarial detection with model interpretation. In SIGKDD . 1803–1811
Ninghao Liu, Hongxia Yang, and Xia Hu. 2018 · 2018
Earlier work this paper cites.
Decoupled Weight Decay Regularization. In ICLR
Ilya Loshchilov and Frank Hutter. 2018 · 2018
Earlier work this paper cites.
Visual interpretability for deep learning: a survey
Quan-shi Zhang and Song-Chun Zhu. 2018 · 2018
Earlier work this paper cites.
Techniques for interpretable machine learning
Mengnan Du, Ninghao Liu, and Xia Hu. 2019a · 2019
Earlier work this paper cites.
Towards automatic concept-based explanations
Amirata Ghorbani, James Wexler, James Y Zou, and Been Kim. 2019 · 2019
Earlier work this paper cites.
Designing and Interpreting Probes with Control Tasks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) . 2733–2743
John Hewitt and Percy Liang. 2019 · 2019
Earlier work this paper cites.
A structural probe for finding syntax in word representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . 4129–4138
John Hewitt and Christopher D Manning. 2019 · 2019
Earlier work this paper cites.
Linguistic knowledge and transferability of contextual representations
Nelson F Liu, Matt Gardner, Yonatan Belinkov, Matthew E Peters, and Noah A Smith. 2019 · 2019
Earlier work this paper cites.
Layer-wise relevance propagation: an overview
Grégoire Montavon, Alexander Binder, Sebastian Lapuschkin, Wojciech Samek, and Klaus-Robert Müller. 2019 · 2019
Earlier work this paper cites.
Definitions, methods, and applications in interpretable machine learning
W James Murdoch, Chandan Singh, Karl Kumbier, Reza Abbasi-Asl, and Bin Yu. 2019 · 2019
Earlier work this paper cites.
Language Models as Knowledge Bases?. In EMNLP-IJCNLP . 2463–2473
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Explain Yourself! Leveraging Language Models for Commonsense Reasoning. In Proceedings of the 57th ACL . 4932–4942
Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, and Richard Socher. 2019 · 2019
Earlier work this paper cites.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Cynthia Rudin. 2019 · 2019
Earlier work this paper cites.
Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned. In Proceedings of the 57th ACL . 5797–5808
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Earlier work this paper cites.
Learning from Explanations with Neural Execution Tree. In ICLR
Ziqi Wang, Yujia Qin, Wenxuan Zhou, Jun Yan, Qinyuan Ye, Leonardo Neves, Zhiyuan Liu, and Xiang Ren. 2019 · 2019
Earlier work this paper cites.
Evaluating explanation without ground truth in interpretable machine learning
Fan Yang, Mengnan Du, and Xia Hu. 2019 · 2019
Earlier work this paper cites.
Thread: circuits
Nick Cammarata, Shan Carter, Gabriel Goh, Chris Olah, Michael Petrov, Ludwig Schubert, Chelsea Voss, Ben Egan, and Swee Kiat Lim. 2020 · 2020
Earlier work this paper cites.
Analyzing Redundancy in Pretrained Transformer Models. In EMNLP
Fahim Dalvi, Hassan Sajjad, Nadir Durrani, and Yonatan Belinkov. 2020 · 2020
Earlier work this paper cites.
A Survey of the State of Explainable AI for Natural Language Processing. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of ACL and the 10th IJCNLP . 447–459
Marina Danilevsky, Kun Qian, Ranit Aharonov, Yannis Katsis, Ban Kawas, and Prithviraj Sen. 2020 · 2020
Earlier work this paper cites.
ERASER: A Benchmark to Evaluate Rationalized NLP Models. In ACL . 4443–4458
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C Wallace. 2020 · 2020
Earlier work this paper cites.
Shortcut learning in deep neural networks
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. 2020 · 2020
Earlier work this paper cites.
Retrieval augmented language model pre-training. In ICML
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020 · 2020
Earlier work this paper cites.
Explaining Black Box Predictions and Unveiling Data Artifacts through Influence Functions. In ACL . 5553–5563
Xiaochuang Han, Byron C Wallace, and Yulia Tsvetkov. 2020 · 2020
Earlier work this paper cites.
Leveraging passage retrieval with generative models for open domain question answering
Gautier Izacard and Edouard Grave. 2020 · 2020
Earlier work this paper cites.
Concept bottleneck models. In ICML . PMLR, 5338–5348
Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pierson, Been Kim, and Percy Liang. 2020 · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
Generate neural template explanations for recommendation. In CIKM . 755–764
Lei Li, Yongfeng Zhang, and Li Chen. 2020 · 2020
Earlier work this paper cites.
Interpreting GPT: the Logit Lens
nostalgebraist. 2020 · 2020
Earlier work this paper cites.
How context affects language models’ factual predictions
Fabio Petroni, Patrick Lewis, Aleksandra Piktus, Tim Rocktäschel, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel. 2020 · 2020
Earlier work this paper cites.
Information-Theoretic Probing for Linguistic Structure. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . 4609–4622
Tiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod, Adina Williams, and Ryan Cotterell. 2020 · 2020
Earlier work this paper cites.
Estimating training data influence by tracing gradient descent
G. Pruthi, F. Liu, S. Kale, and M. Sundararajan. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Earlier work this paper cites.
A survey on explainable artificial intelligence (xai): Toward medical xai
Erico Tjoa and Cuntai Guan. 2020 · 2020
Earlier work this paper cites.
Analyzing the source and target contributions to predictions in neural machine translation
E. Voita, R. Sennrich, and I. Titov. 2020 · 2020
Earlier work this paper cites.
Information-Theoretic Probing with Minimum Description Length. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) . 183–196
Elena Voita and Ivan Titov. 2020 · 2020
Earlier work this paper cites.
Fact or Fiction: Verifying Scientific Claims. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) . 7534–7550
David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020 · 2020
Earlier work this paper cites.
Improving VQA and its Explanations by Comparing Competing Explanations
Jialin Wu, Liyan Chen, and Raymond J Mooney. 2020a · 2020
Earlier work this paper cites.
Learning diverse and discriminative representations via the principle of maximal coding rate reduction
Yaodong Yu, Kwan Ho Ryan Chan, Chong You, Chaobing Song, and Yi Ma. 2020 · 2020
Earlier work this paper cites.
Neural additive models: Interpretable machine learning with neural nets
Rishabh Agarwal, Levi Melnick, Nicholas Frosst, Xuezhou Zhang, Ben Lengerich, Rich Caruana, and Geoffrey E Hinton. 2021 · 2021
Earlier work this paper cites.
Rationalization through Concepts
Diego Matteo Antognini and Boi Faltings. 2021 · 2021
Earlier work this paper cites.
Rationale-inspired natural language explanations with commonsense
Bodhisattwa Prasad Majumder1 Oana-Maria Camburu, Thomas Lukasiewicz, and Julian McAuley. 2021 · 2021
Earlier work this paper cites.
What to learn, and how: Toward effective learning from rationales
Samuel Carton, Surya Kanoria, and Chenhao Tan. 2021 · 2021
Earlier work this paper cites.
Probing BERT in Hyperbolic Spaces. In International Conference on Learning Representations
Boli Chen, Yao Fu, Guangwei Xu, Pengjun Xie, Chuanqi Tan, Mosha Chen, and Liping Jing. 2021 · 2021
Earlier work this paper cites.
Towards interpreting and mitigating shortcut learning behavior of NLU models
Mengnan Du, Varun Manjunatha, Rajiv Jain, Ruchi Deshpande, Franck Dernoncourt, Jiuxiang Gu, Tong Sun, and Xia Hu. 2021 · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al · 2021
Earlier work this paper cites.
Causal abstractions of neural networks
Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts. 2021 · 2021
Earlier work this paper cites.
Transformer Feed-Forward Layers Are Key-Value Memories. In EMNLP
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021 · 2021
Earlier work this paper cites.
FastIF: Scalable Influence Functions for Efficient Model Interpretation and Debugging. In EMNLP . 10333–10350
Han Guo, Nazneen Rajani, Peter Hase, Mohit Bansal, and Caiming Xiong. 2021 · 2021
Earlier work this paper cites.
BERT meets shapley: Extending SHAP explanations to transformer-based classifiers. In Proceedings of the EACL Hackashop on News Media Content Analysis and Automated Report Generation . 16–21
Enja Kokalj, Blaž Škrlj, Nada Lavrač, Senja Pollak, and Marko Robnik-Šikonja. 2021 · 2021
Earlier work this paper cites.
Adversarial attacks and defenses: An interpretation perspective
Ninghao Liu, Mengnan Du, Ruocheng Guo, Huan Liu, and Xia Hu. 2021 · 2021
Earlier work this paper cites.
Exploring the Role of BERT Token Representations to Explain Sentence Probing Results. In EMNLP . 792–806
Hosein Mohebbi, Ali Modarressi, and Mohammad Taher Pilehvar. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision. In ICML . PMLR, 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Retrieval augmentation reduces hallucination in conversation
K. Shuster, S. Poff, M. Chen, D. Kiela, and J. Weston. 2021 · 2021
Earlier work this paper cites.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki. 2021 · 2021
Earlier work this paper cites.
Robustness to spurious correlations in text classification via automatically generated counterfactuals. In AAAI , Vol. 35. 14024–14031
Zhao Wang and Aron Culotta. 2021 · 2021
Earlier work this paper cites.
Analysing bias in spoken language assessment using concept activation vectors. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 7753–7757
Xizi Wei, Mark JF Gales, and Kate M Knill. 2021 · 2021
Earlier work this paper cites.
On explaining your explanations of bert: An empirical study with sequence classification
Zhengxuan Wu and Desmond C Ong. 2021 · 2021
Earlier work this paper cites.
Lirex: Augmenting language inference with relevant explanations. In AAAI , Vol. 35. 14532–14539
Xinyan Zhao and VG Vinod Vydiswaran. 2021 · 2021
Earlier work this paper cites.
Interpretable ranking with generalized additive models. In WSDM
Honglei Zhuang, Xuanhui Wang, Michael Bendersky, Alexander Grushetsky, Yonghui Wu, Petr Mitrichev, Ethan Sterling, Nathan Bell, Walker Ravina, and Hai Qian. 2021 · 2021
Earlier work this paper cites.
Towards Tracing Knowledge in Language Models Back to the Training Data. In Findings of EMNLP . 2429–2446
Ekin Akyurek, Tolga Bolukbasi, Frederick Liu, Binbin Xiong, Ian Tenney, Jacob Andreas, and Kelvin Guu. 2022 · 2022
Earlier work this paper cites.
Concept gradient: Concept-based interpretation without linear assumption
Andrew Bai, Chih-Kuan Yeh, Pradeep Ravikumar, Neil YC Lin, and Cho-Jui Hsieh. 2022 · 2022
Earlier work this paper cites.
Attributed question answering: Evaluation and modeling for attributed large language models
Bernd Bohnet, Vinh Q Tran, Pat Verga, Roee Aharoni, Daniel Andor, Livio Baldini Soares, Massimiliano Ciaramita, Jacob Eisenstein, Kuzman Ganchev, Jonathan Herzig, et al · 2022
Earlier work this paper cites.
Discovering latent knowledge in language models without supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. 2022 · 2022
Earlier work this paper cites.
M6-rec: Generative pretrained language models are open-ended recommender systems
Zeyu Cui, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. 2022 · 2022
Earlier work this paper cites.
Knowledge Neurons in Pretrained Transformers. In ACL
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2022 · 2022
Earlier work this paper cites.
Is gpt-3 a good data annotator?
B. Ding, C. Qin, L. Liu, Y. K. Chia, S. Joty, B. Li, and L. Bing. 2022 · 2022
Earlier work this paper cites.
Recommendation as language processing (rlp): A unified pretrain, personalized prompt & predict paradigm (p5). In RecSys . 299–315
Shijie Geng, Shuchang Liu, Zuohui Fu, Yingqiang Ge, and Yongfeng Zhang. 2022 · 2022
Earlier work this paper cites.
Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing . 30–45
Mor Geva, Avi Caciularu, Kevin Wang, and Yoav Goldberg. 2022 · 2022
Earlier work this paper cites.
Xiaochuang Han and Yulia Tsvetkov. 2022 · 2022
Earlier work this paper cites.
Large language models are reasoning teachers
Namgyu Ho, Laura Schmid, and Se-Young Yun. 2022 · 2022
Earlier work this paper cites.
Introduction to Transformers for NLP: With the Hugging Face Library and Models to Solve Problems
Shashank Mohan Jain. 2022 · 2022
Earlier work this paper cites.
Decomposed Prompting: A Modular Approach for Solving Complex Tasks. In The Eleventh ICLR
Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal. 2022 · 2022
Earlier work this paper cites.
Large Language Models are Zero-Shot Reasoners. In Advances in Neural Information Processing Systems , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35. Curran Associates, Inc., 22199–22213
Takeshi Kojima, Shixiang (Shane) Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
Earlier work this paper cites.
Large language models with controllable working memory
Daliang Li, Ankit Singh Rawat, Manzil Zaheer, Xin Wang, Michal Lukasik, Andreas Veit, Felix Yu, and Sanjiv Kumar. 2022b · 2022
Earlier work this paper cites.
Explanations from large language models make small reasoners better
Shiyang Li, Jianshu Chen, Yelong Shen, Zhiyu Chen, Xinlu Zhang, Zekun Li, Hong Wang, Jing Qian, Baolin Peng, Yi Mao, et al · 2022
Earlier work this paper cites.
Wei Li, Wenhao Wu, Moye Chen, Jiachen Liu, Xinyan Xiao, and Hua Wu. 2022c · 2022
Cited alongside, same era.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022 · 2022
Cited alongside, same era.
Teaching small language models to reason
L. C. Magister, J. Mallinson, J. Adamek, E. Malmi, and A. Severyn. 2022 · 2022
Cited alongside, same era.
Locating and editing factual associations in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022a · 2022
Cited alongside, same era.
Visual Classification via Description from Large Language Models. In The Eleventh ICLR
Sachit Menon and Carl Vondrick. 2022 · 2022
Cited alongside, same era.
Beyond Labels: Empowering Human Annotators with Natural Language Explanations through a Novel Active-Learning Architecture. In Findings of the Association for Computational Linguistics: EMNLP 2023 . 11629–11643
Bingsheng Yao, Ishan Jindal, Lucian Popa, Yannis Katsis, Sayan Ghosh, Lihong He, Yuxuan Lu, Shashank Srivastava, Yunyao Li, James Hendler, et al · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L Griffiths, Yuan Cao, and Karthik Narasimhan. 2023b · 2023
Later among the works it cites.
Making retrieval-augmented language models robust to irrelevant context
O. Yoran, T. Wolfson, O. Ram, and J. Berant. 2023 · 2023
Later among the works it cites.
Unlearning bias in language models by partitioning gradients. In Findings of the Association for Computational Linguistics: ACL 2023 . 6032–6048
Charles Yu, Sullam Jeoung, Anish Kasi, Pengfei Yu, and Heng Ji. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
How to dissect a Muppet: The structure of transformer embedding spaces
Timothee Mickus, Denis Paperno, and Mathieu Constant. 2022 · 2022
Cited alongside, same era.
The singular value decompositions of transformer weight matrices are highly interpretable
Beren Millidge and Sid Black. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Interpretable machine learning: Fundamental principles and 10 grand challenges
Cynthia Rudin, Chaofan Chen, Zhi Chen, Haiyang Huang, Lesia Semenova, and Chudi Zhong. 2022 · 2022
Cited alongside, same era.
Polysemanticity and capacity in neural networks
A. Scherlis, K. Sachan, A. S Jermyn, J. Benton, and B. Shlegeris. 2022 · 2022
Cited alongside, same era.
Scaling up influence functions. In AAAI , Vol. 36. 8179–8186
Andrea Schioppa, Polina Zablotskaia, David Vilar, and Artem Sokolov. 2022 · 2022
Cited alongside, same era.
Finding Skill Neurons in Pre-trained Transformer-based Language Models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing . 11132–11152
Xiaozhi Wang, Kaiyue Wen, Zhengyan Zhang, Lei Hou, Zhiyuan Liu, and Juanzi Li. 2022b · 2022
Cited alongside, same era.
Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, and Alex Smola. 2023 · 2023
Later among the works it cites.
Automated Natural Language Explanation of Deep Visual Neurons with Large Models
Chenxu Zhao, Wei Qian, Yucheng Shi, Mengdi Huai, and Ninghao Liu. 2023b · 2023
Later among the works it cites.
Explainability for LLMs: A survey
H. Zhao, H. Chen, F. Yang, N. Liu, H. Deng, H. Cai, S. Wang, D. Yin, and M. Du. 2023a · 2023
Later among the works it cites.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Later among the works it cites.
Interpretation of Time-Series Deep Models: A Survey
Z. Zhao, Y. Shi, S. Wu, F. Yang, W. Song, and N. Liu. 2023c · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Later among the works it cites.
MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions
Zexuan Zhong, Zhengxuan Wu, Christopher D Manning, Christopher Potts, and Danqi Chen. 2023 · 2023
Later among the works it cites.
Yaochen Zhu, Jing Ma, and Jundong Li. 2023 · 2023
Later among the works it cites.
Representation Engineering: A Top-Down Approach to AI Transparency
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Z. Wang, J Z. Kolter, and M. Fredrikson. 2023b · 2023
Later among the works it cites.
Do Language Models Know When They’re Hallucinating References?. In Findings of the Association for Computational Linguistics: EACL 2024 . 912–928
Ayush Agrawal, Mirac Suzgun, Lester Mackey, and Adam Kalai. 2024 · 2024
Closest in time.
Distinguishing the Knowable from the Unknowable with Language Models
Gustaf Ahdritz, Tian Qin, Nikhil Vyas, Boaz Barak, and Benjamin L Edelman. 2024 · 2024
Closest in time.
Special characters attack: Toward scalable training data extraction from large language models
Yang Bai, Ge Pei, Jindong Gu, Yong Yang, and Xingjun Ma. 2024 · 2024
Closest in time.
Towards Inference-time Category-wise Safety Steering for Large Language Models. In Neurips Safe Generative AI Workshop 2024
Amrita Bhattacharjee, Shaona Ghosh, Traian Rebedea, and Christopher Parisien. 2024 · 2024
Closest in time.
Using Dictionary Learning Features as Classifiers
Trenton Bricken, Jonathan Marcus, Siddharth Mishra-Sharma, Meg Tong, Ethan Perez, Mrinank Sharma, Kelley Rivoire, and Thomas Henighan. 2024 · 2024
Closest in time.
SelfIE: Self-Interpretation of Large Language Model Embeddings
Haozhe Chen, Carl Vondrick, and Chengzhi Mao. 2024c · 2024
Closest in time.
Unlocking the capabilities of thought: A reasoning boundary framework to quantify and optimize chain-of-thought
Qiguang Chen, Libo Qin, Jiaqi Wang, Jingxuan Zhou, and Wanxiang Che. 2024a · 2024
Closest in time.
ChainLM: Empowering Large Language Models with Improved Chain-of-Thought Prompting. In LREC/COLING
Xiaoxue Cheng, Junyi Li, Wayne Xin Zhao, and Ji-Rong Wen. 2024 · 2024
Closest in time.
Jump to Conclusions: Short-Cutting Transformers with Linear Transformations. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) . 9615–9625
Alexander Yom Din, Taelin Karidi, Leshem Choshen, and Mor Geva. 2024 · 2024
Closest in time.
Do LLMs Know about Hallucination? An Empirical Investigation of LLM’s Hidden States
H. Duan, Yi Yang, and K. Y. Tam. 2024 · 2024
Closest in time.
Clément Dumas, Chris Wendler, Veniamin Veselovsky, Giovanni Monea, and Robert West. 2024 · 2024
Closest in time.
Evaluating Feature Steering: A Case Study in Mitigating Social Biases
Esin Durmus, Alex Tamkin, Jack Clark, Jerry Wei, Jonathan Marcus, Joshua Batson, Kunal Handa, Liane Lovitt, Meg Tong, Miles McCain, Oliver Rausch, Saffron Huang, Sam Bowman, Stuart Ritchie, Tom Henighan, and Deep Ganguli. 2024 · 2024
Closest in time.
Not All Layers of LLMs Are Necessary During Inference
Siqi Fan, Xin Jiang, Xiang Li, Xuying Meng, Peng Han, Shuo Shang, Aixin Sun, Yequan Wang, and Zhongyuan Wang. 2024 · 2024
Closest in time.
A Primer on the Inner Workings of Transformer-based Language Models
Javier Ferrando, Gabriele Sarti, Arianna Bisazza, and Marta R Costa-jussà. 2024 · 2024
Closest in time.
Think or Remember? Detecting and Directing LLMs Towards Memorization or Generalization
Yi-Fu Fu, Yu-Chieh Tu, Tzu-Ling Cheng, Cheng-Yu Lin, Yi-Ting Yang, Heng-Yi Liu, Keng-Te Liao, Da-Cheng Juan, and Shou-De Lin. 2024 · 2024
Closest in time.
Scaling and evaluating sparse autoencoders
Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu. 2024 · 2024
Closest in time.
Yuan Ge, Yilun Liu, Chi Hu, Weibin Meng, Shimin Tao, Xiaofeng Zhao, Hongxia Ma, Li Zhang, Boxing Chen, Hao Yang, Bei Li, Tong Xiao, and Jingbo Zhu. 2024 · 2024
Closest in time.
A Unified Framework for Model Editing
Akshat Gupta, Dev Sajnani, and Gopala Anumanchipalli. 2024 · 2024
Closest in time.
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms. In ICML 2024 Workshop on Mechanistic Interpretability
Michael Hanna, Sandro Pezzelle, and Yonatan Belinkov. [n. d.] · 2024
Closest in time.
Does localization inform editing? surprising differences in causality-based localization vs. knowledge editing in language models
Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun. 2024 · 2024
Closest in time.
Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence
Jinghan He, Kuan Zhu, Haiyun Guo, Junfeng Fang, Zhenglin Hua, Yuheng Jia, Ming Tang, Tat-Seng Chua, and Jinqiao Wang. 2024 · 2024
Closest in time.
Do LLMs" know" internally when they follow instructions?
Juyeon Heo, Christina Heinze-Deml, Oussama Elachqar, Shirley Ren, Udhay Nallasamy, Andy Miller, Kwan Ho Ryan Chan, and Jaya Narain. 2024 · 2024
Closest in time.
Steering LLMs’ Behavior with Concept Activation Vectors
Ruixuan Huang. 2024 · 2024
Closest in time.
Dishonesty in helpful and harmless alignment
Youcheng Huang, Jingkun Tang, Duanyu Feng, Zheng Zhang, Wenqiang Lei, Jiancheng Lv, and Anthony G Cohn. 2024 · 2024
Closest in time.
Mindstar: Enhancing math reasoning in pre-trained llms at inference time
Jikun Kang, Xin Zhe Li, Xi Chen, Amirreza Kazemi, Qianyi Sun, Boxing Chen, Dong Li, Xu He, Quan He, Feng Wen, et al · 2024
Closest in time.
Post hoc explanations of language models can improve language models
Satyapriya Krishna, Jiaqi Ma, Dylan Slack, Asma Ghandeharioun, Sameer Singh, and Himabindu Lakkaraju. 2024 · 2024
Closest in time.
Fine-tuning chatgpt for automatic scoring
Ehsan Latif and Xiaoming Zhai. 2024 · 2024
Closest in time.
A mechanistic understanding of alignment algorithms: A case study on dpo and toxicity
Andrew Lee, Xiaoyan Bai, Itamar Pres, Martin Wattenberg, Jonathan K Kummerfeld, and Rada Mihalcea. 2024 · 2024
Closest in time.
Inference-time intervention: Eliciting truthful answers from a language model
Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. 2024b · 2024
Closest in time.
Open the Pandora’s Box of LLMs: Jailbreaking LLMs through Representation Engineering
Tianlong Li, Xiaoqing Zheng, and Xuanjing Huang. 2024d · 2024
Closest in time.
On the intrinsic self-correction capability of LLMs: Uncertainty and latent concept
Guangliang Liu, Haitao Mao, Bochuan Cao, Zhiyu Xue, Xitong Zhang, Rongrong Wang, Jiliang Tang, and Kristen Johnson. 2024b · 2024
Closest in time.
Lost in the middle: How language models use long contexts
Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024a · 2024
Closest in time.
From Understanding to Utilization: A Survey on Explainability for Large Language Models
Haoyan Luo and Lucia Specia. 2024 · 2024
Closest in time.
Source2synth: Synthetic data generation and curation grounded in real data sources
Alisia Lupidi, Carlos Gemmell, Nicola Cancedda, Jane Dwivedi-Yu, Jason Weston, Jakob Foerster, Roberta Raileanu, and Maria Lomeli. 2024 · 2024
Closest in time.
Simple probes can catch sleeper agents
Monte MacDiarmid, Timothy Maxwell, Nicholas Schiefer, Jesse Mu, Jared Kaplan, David Duvenaud, Sam Bowman, Alex Tamkin, Ethan Perez, Mrinank Sharma, et al · 2024
Closest in time.
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
Samuel Marks, Can Rager, Eric J Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller. 2024 · 2024
Closest in time.
ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
Xin Men, Mingyu Xu, Qingyu Zhang, Bingning Wang, Hongyu Lin, Yaojie Lu, Xianpei Han, and Weipeng Chen. 2024 · 2024
Closest in time.
TextCAVs: Debugging vision models using text. In International Conference on Medical Image Computing and Computer-Assisted Intervention . Springer, 99–109
Angus Nicolson, Yarin Gal, and J Alison Noble. 2024 · 2024
Closest in time.
LatentQA: Teaching LLMs to Decode Activations Into Natural Language
Alexander Pan, Lijie Chen, and Jacob Steinhardt. 2024 · 2024
Closest in time.
Iterative reasoning preference optimization
Richard Yuanzhe Pang, Weizhe Yuan, He He, Kyunghyun Cho, Sainbayar Sukhbaatar, and Jason Weston. 2024 · 2024
Closest in time.
Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented Generation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, 6037–6053
Jirui Qi, Gabriele Sarti, Raquel Fernández, and Arianna Bisazza. 2024 · 2024
Closest in time.
Chen Qian, Dongrui Liu, Jie Zhang, Yong Liu, and Jing Shao. 2024 · 2024
Closest in time.
From Prejudice to Parity: A New Approach to Debiasing Large Language Model Word Embeddings
Aishik Rakshit, Smriti Singh, Shuvam Keshari, Arijit Ghosh Chowdhury, Vinija Jain, and Aman Chadha. 2024 · 2024
Closest in time.
Steering Llama 2 via Contrastive Activation Addition. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 15504–15522
Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Turner. 2024 · 2024
Closest in time.
FIND: A Function Description Benchmark for Evaluating Interpretability Methods
Sarah Schwettmann, Tamar Shaham, Joanna Materzynska, Neil Chowdhury, Shuang Li, Jacob Andreas, David Bau, and Antonio Torralba. 2024 · 2024
Closest in time.
Incremental Residual Concept Bottleneck Models
Chenming Shang, Shiji Zhou, Yujiu Yang, Hengyuan Zhang, Xinzhe Ni, and Yuwang Wang. 2024 · 2024
Closest in time.
Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2024 · 2024
Closest in time.
Mimic: Minimally modified counterfactuals in the representation space
Shashwat Singh, Shauli Ravfogel, Jonathan Herzig, Roee Aharoni, Ryan Cotterell, and Ponnurangam Kumaraguru. 2024 · 2024
Closest in time.
Improving instruction-following in language models through activation steering
Alessandro Stolfo, Vidhisha Balachandran, Safoora Yousefi, Eric Horvitz, and Besmira Nushi. 2024 · 2024
Closest in time.
Large Language Models for Data Annotation: A Survey
Zhen Tan, Alimohammad Beigi, Song Wang, Ruocheng Guo, Amrita Bhattacharjee, Bohan Jiang, Mansooreh Karami, Jundong Li, Lu Cheng, and Huan Liu. 2024 · 2024
Closest in time.
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan. 2024 · 2024
Closest in time.
Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. 2024 · 2024
Closest in time.
Dan is my new friend
walkerspider. 2022 · 2024
Closest in time.
Data advisor: Dynamic data curation for safety alignment of large language models
Fei Wang, Ninareh Mehrabi, Palash Goyal, Rahul Gupta, Kai-Wei Chang, and Aram Galstyan. 2024b · 2024
Closest in time.
Detoxifying large language models via knowledge editing
Mengru Wang, Ningyu Zhang, Ziwen Xu, Zekun Xi, Shumin Deng, Yunzhi Yao, Qishen Zhang, Linyi Yang, Jindong Wang, and Huajun Chen. 2024d · 2024
Closest in time.
Chain-of-Thought Reasoning Without Prompting
Xuezhi Wang and Denny Zhou. 2024 · 2024
Closest in time.
Describe, explain, plan and select: interactive planning with LLMs enables open-world multi-task agents
Zihao Wang, Shaofei Cai, Guanzhou Chen, Anji Liu, Xiaojian Shawn Ma, and Yitao Liang. 2024a · 2024
Closest in time.
Hui Wei, Shenghua He, Tian Xia, Fei Liu, Andy Wong, Jingyang Lin, and Mei Han. 2024b · 2024
Closest in time.
“According to…”: Prompting Language Models Improves Quoting from Pre-Training Data. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) . 2288–2301
Orion Weller, Marc Marone, Nathaniel Weir, Dawn Lawrie, Daniel Khashabi, and Benjamin Van Durme. 2024 · 2024
Closest in time.
How faithful are RAG models? Quantifying the tug-of-war between RAG and LLMs’ internal prior
K. Wu, E. Wu, and J. Zou. 2024b · 2024
Closest in time.
Thinking llms: General instruction following with thought generation
Tianhao Wu, Janice Lan, Weizhe Yuan, Jiantao Jiao, Jason Weston, and Sainbayar Sukhbaatar. 2024a · 2024
Closest in time.
Hallucination is inevitable: An innate limitation of large language models
Ziwei Xu, Sanjay Jain, and Mohan Kankanhalli. 2024 · 2024
Closest in time.
Yuqing Yang, Yan Ma, and Pengfei Liu. 2024 · 2024
Closest in time.
Xunjian Yin, Xu Zhang, Jie Ruan, and Xiaojun Wan. 2024b · 2024
Closest in time.
Mechanistic understanding and mitigation of language model non-factual hallucinations
Lei Yu, Meng Cao, Jackie Chi Kit Cheung, and Yue Dong. 2024a · 2024
Closest in time.
Robust LLM safeguarding via refusal feature adversarial training
Lei Yu, Virginie Do, Karen Hambardzumyan, and Nicola Cancedda. 2024b · 2024
Closest in time.
Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking
Eric Zelikman, Georges Harik, Yijia Shao, Varuna Jayasiri, Nick Haber, and Noah D Goodman. 2024 · 2024
Closest in time.
Mitigating the privacy issues in retrieval-augmented generation (rag) via pure synthetic data
Shenglai Zeng, Jiankun Zhang, Pengfei He, Jie Ren, Tianqi Zheng, Hanqing Lu, Han Xu, Hui Liu, Yue Xing, and Jiliang Tang. 2024b · 2024
Closest in time.
Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi. 2024a · 2024
Closest in time.
Chain of preference optimization: Improving chain-of-thought reasoning in llms
Xuan Zhang, Chao Du, Tianyu Pang, Qian Liu, Wei Gao, and Min Lin. 2024a · 2024
Closest in time.
Seeing clearly by layer two: Enhancing attention heads to alleviate hallucination in lvlms
Xiaofeng Zhang, Yihao Quan, Chaochen Gu, Chen Shen, Xiaosong Yuan, Shaotian Yan, Hao Cheng, Kaijie Wu, and Jieping Ye. 2024c · 2024
Closest in time.
Learn Beyond The Answer: Training Language Models with Reflection for Mathematical Reasoning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing . 14720–14738
Zhihan Zhang, Tao Ge, Zhenwen Liang, Wenhao Yu, Dian Yu, Mengzhao Jia, Dong Yu, and Meng Jiang. 2024b · 2024
Closest in time.
Opening the Black Box of LLMs: Two Views on Holistic Interpretability
H. Zhao, F. Yang, H. Lakkaraju, and M. Du. 2024 · 2024
Closest in time.
On Prompt-Driven Safeguarding for Large Language Models. In ICLR 2024 Workshop on Secure and Trustworthy Large Language Models
Chujie Zheng, Fan Yin, Hao Zhou, Fandong Meng, Jie Zhou, Kai-Wei Chang, Minlie Huang, and Nanyun Peng. [n. d.] · 2024
Closest in time.
On Prompt-Driven Safeguarding for Large Language Models. In Forty-first International Conference on Machine Learning
Chujie Zheng, Fan Yin, Hao Zhou, Fandong Meng, Jie Zhou, Kai-Wei Chang, Minlie Huang, and Nanyun Peng. 2024 · 2024
Closest in time.
DETAIL: Task DEmonsTration Attribution for Interpretable In-context Learning
Zijian Zhou, Xiaoqiang Lin, Xinyi Xu, Alok Prakash, Daniela Rus, and Bryan Kian Hsiang Low. 2024 · 2024
Closest in time.
Collaborative large language model for recommender systems. In WWW
Yaochen Zhu, Liang Wu, Qi Guo, Liangjie Hong, and Jundong Li. 2024 · 2024
Closest in time.
Refusal in language models is mediated by a single direction
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda. 2025 · 2025
Closest in time.
The Unreasonable Ineffectiveness of the Deeper Layers
Andrey Gromov, Kushal Tirumala, Hassan Shapourian, Paolo Glorioso, and Daniel A. Roberts. 2025 · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Enhancing Automated Interpretability with Output-Centric Feature Descriptions
Yoav Gur-Arieh, Roy Mayan, Chen Agassy, Atticus Geiger, and Mor Geva. 2025 · 2025
Closest in time.
Zirui He, Haiyan Zhao, Yiran Qiao, Fan Yang, Ali Payani, Jing Ma, and Mengnan Du. 2025 · 2025
Closest in time.
Sparse Autoencoders Can Interpret Randomly Initialized Transformers
Thomas Heap, Tim Lawson, Lucy Farnik, and Laurence Aitchison. 2025 · 2025
Closest in time.
Thinkprune: Pruning long chain-of-thought of llms via reinforcement learning
Bairu Hou, Yang Zhang, Jiabao Ji, Yujian Liu, Kaizhi Qian, Jacob Andreas, and Shiyu Chang. 2025 · 2025
Closest in time.
O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning
Haotian Luo, Li Shen, Haiying He, Yibo Wang, Shiwei Liu, Wei Li, Naiqiang Tan, Xiaochun Cao, and Dacheng Tao. 2025 · 2025
Closest in time.
How to Steer LLM Latents for Hallucination Detection?
Seongheon Park, Xuefeng Du, Min-Hsuan Yeh, Haobo Wang, and Yixuan Li. 2025 · 2025
Closest in time.
Sparse Autoencoders Trained on the Same Data Learn Different Features
Gonçalo Paulo and Nora Belrose. 2025 · 2025
Closest in time.
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
Swarnadeep Saha, Xian Li, Marjan Ghazvininejad, Jason Weston, and Tianlu Wang. 2025 · 2025
Closest in time.
SAKE: Steering Activations for Knowledge Editing
Marco Scialanga, Thibault Laugel, Vincent Grari, and Marcin Detyniecki. 2025 · 2025
Closest in time.
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
Dong Shu, Xuansheng Wu, Haiyan Zhao, Daking Rai, Ziyu Yao, Ninghao Liu, and Mengnan Du. 2025 · 2025
Closest in time.
Mitigating Memorization in LLMs using Activation Steering
Manan Suri, Nishit Anand, and Amisha Bhaskar. 2025 · 2025
Closest in time.
Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders
Xuansheng Wu, Jiayi Yuan, Wenlin Yao, Xiaoming Zhai, and Ninghao Liu. 2025c · 2025
Closest in time.
AXBENCH: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
Zhengxuan Wu, Aryaman Arora, Atticus Geiger, Zheng Wang, Jing Huang, Dan Jurafsky, Christopher D Manning, and Christopher Potts. 2025a · 2025
Closest in time.
Think When You Need: Self-Adaptive Chain-of-Thought Learning
Junjie Yang, Ke Lin, and Xing Yu. 2025a · 2025
Closest in time.
Representation Bending for Large Language Model Safety
Ashkan Yousefpour, Taeheon Kim, Ryan S Kwon, Seungbeen Lee, Wonje Jeung, Seungju Han, Alvin Wan, Harrison Ngan, Youngjae Yu, and Jonghyun Choi. 2025 · 2025
Closest in time.