Fetching the paper…
Reading the bibliography…
The rise of foundation models has transformed machine learning research, prompting efforts to uncover their inner workings and develop more efficient and reliable applications for better control.
Visualizing attention in transformer-based language models
Jesse Vig. 2019 · 1904
Earlier work this paper cites.
Learning to deceive with attention-based explanations
Danish Pruthi, Mansi Gupta, Bhuwan Dhingra, Graham Neubig, and Zachary C Lipton. 2019 · 1909
Earlier work this paper cites.
Attention interpretability across nlp tasks
Shikhar Vashishth, Shyam Upadhyay, Gaurav Singh Tomar, and Manaal Faruqui. 2019 · 1909
Earlier work this paper cites.
Direct and indirect effects
Judea Pearl. 2001 · 2001
Earlier work this paper cites.
Differential privacy
Cynthia Dwork. 2006 · 2006
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2006
Earlier work this paper cites.
Explaining classifications for individual instances
Marko Robnik-Šikonja and Igor Kononenko. 2008 · 2008
Earlier work this paper cites.
Analyzing individual neurons in pre-trained language models
Nadir Durrani, Hassan Sajjad, Fahim Dalvi, and Yonatan Belinkov. 2020 · 2010
Earlier work this paper cites.
Interpretation and identification of causal mediation
Judea Pearl. 2014 · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015 · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2015 · 2015
Earlier work this paper cites.
" why should i trust you?" explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
David Bau, Bolei Zhou, Aditya Khosla, Aude Oliva, and Antonio Torralba. 2017 · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee. 2017 · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. 2017 · 2017
Earlier work this paper cites.
Visualizing deep neural network decisions: Prediction difference analysis
Luisa M Zintgraf, Taco S Cohen, Tameem Adel, and Max Welling. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin. 2018 · 2018
Earlier work this paper cites.
Multimodal explanations: Justifying decisions and pointing to the evidence
Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata, Anna Rohrbach, Bernt Schiele, Trevor Darrell, and Marcus Rohrbach. 2018 · 2018
Earlier work this paper cites.
Deep learning using rectified linear units (relu)
Abien Fred Agarap. 2019 · 2019
Earlier work this paper cites.
What is one grain of sand in the desert? analyzing individual neurons in deep nlp models
Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, Anthony Bau, and James Glass. 2019 · 2019
Earlier work this paper cites.
Attention is not Explanation
Sarthak Jain and Byron C. Wallace. 2019 · 2019
Earlier work this paper cites.
What does BERT learn about the structure of language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019 · 2019
Earlier work this paper cites.
Definitions, methods, and applications in interpretable machine learning
W James Murdoch, Chandan Singh, Karl Kumbier, Reza Abbasi-Asl, and Bin Yu. 2019 · 2019
Earlier work this paper cites.
Sanvis: Visual analytics for understanding self-attention networks
Cheonbok Park, Inyoup Na, Yongjang Jo, Sungbok Shin, Jaehyo Yoo, Bum Chul Kwon, Jian Zhao, Hyungjong Noh, Yeonsoo Lee, and Jaegul Choo. 2019 · 2019
Earlier work this paper cites.
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Earlier work this paper cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Earlier work this paper cites.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter. 2019 · 2019
Earlier work this paper cites.
Behind the scene: Revealing the secrets of pre-trained vision-and-language models
Jize Cao, Zhe Gan, Yu Cheng, Licheng Yu, Yen-Chun Chen, and Jingjing Liu. 2020 · 2020
Earlier work this paper cites.
Probing multimodal embeddings for linguistic properties: the visual-semantic case
Adam Dahlgren Lindström, Johanna Björklund, Suna Bensch, and Frank Drewes. 2020 · 2020
Earlier work this paper cites.
exBERT: A Visual Analysis Tool to Explore Learned Representations in Transformer Models
Benjamin Hoover, Hendrik Strobelt, and Sebastian Gehrmann. 2020 · 2020
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang. 2020 · 2020
Earlier work this paper cites.
Federated learning: Challenges, methods, and future directions
Tian Li, Anit Kumar Sahu, Ameet Talwalkar, and Virginia Smith. 2020 · 2020
Earlier work this paper cites.
Estimating training data influence by tracing gradient descent
Garima Pruthi, Frederick Liu, Satyen Kale, and Mukund Sundararajan. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Earlier work this paper cites.
Intrinsic probing through dimension selection
Lucas Torroba Hennigen, Adina Williams, and Ryan Cotterell. 2020 · 2020
Earlier work this paper cites.
DeeBERT: Dynamic early exiting for accelerating BERT inference
Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin. 2020 · 2020
Earlier work this paper cites.
On the pitfalls of analyzing individual neurons in language models
Omer Antverg and Yonatan Belinkov. 2021 · 2021
Earlier work this paper cites.
Influence functions in deep learning are fragile
Samyadeep Basu, Philip Pope, and Soheil Feizi. 2021 · 2021
Earlier work this paper cites.
Transformer interpretability beyond attention visualization
Hila Chefer, Shir Gur, and Lior Wolf. 2021 · 2021
Earlier work this paper cites.
Knowledge neurons in pretrained transformers
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2021 · 2021
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021 · 2021
Earlier work this paper cites.
Multimodal neurons in artificial neural networks
Gabriel Goh, Nick Cammarata, Chelsea Voss, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah. 2021 · 2021
Earlier work this paper cites.
Self-attention attribution: Interpreting information interactions inside transformer
Yaru Hao, Li Dong, Furu Wei, and Ke Xu. 2021 · 2021
Earlier work this paper cites.
Probing image-language transformers for verb understanding
Lisa Anne Hendricks and Aida Nematzadeh. 2021 · 2021
Earlier work this paper cites.
Natural language descriptions of deep visual features
Evan Hernandez, Sarah Schwettmann, David Bau, Teona Bagashvili, Antonio Torralba, and Jacob Andreas. 2021 · 2021
Earlier work this paper cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yunhsuan Sung, Zhen Li, and Tom Duerig. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Earlier work this paper cites.
Multimodal few-shot learning with frozen language models
Maria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami, Oriol Vinyals, and Felix Hill. 2021 · 2021
Earlier work this paper cites.
Filip: Fine-grained interactive language-image pre-training
Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu. 2021 · 2021
Earlier work this paper cites.
Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors
Zeyu Yun, Yubei Chen, Bruno Olshausen, and Yann LeCun. 2021 · 2021
Earlier work this paper cites.
Vl-interpret: An interactive visualization tool for interpreting vision-language transformers
Estelle Aflalo, Meng Du, Shao-Yen Tseng, Yongfei Liu, Chenfei Wu, Nan Duan, and Vasudev Lal. 2022 · 2022
Earlier work this paper cites.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022 · 2022
Earlier work this paper cites.
Prompt-to-prompt image editing with cross attention control
Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. 2022 · 2022
Earlier work this paper cites.
Diffusers-interpret
João Lages. 2022 · 2022
Earlier work this paper cites.
Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm
Yangguang Li, Feng Liang, Lichen Zhao, Yufeng Cui, Wanli Ouyang, Jing Shao, Fengwei Yu, and Junjie Yan. 2022 · 2022
Earlier work this paper cites.
Multiviz: Towards visualizing and understanding multimodal models
Paul Pu Liang, Yiwei Lyu, Gunjan Chhablani, Nihal Jain, Zihao Deng, Xingbo Wang, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2022 · 2022
Earlier work this paper cites.
Dime: Fine-grained interpretations of multimodal models via disentangled local explanations
Yiwei Lyu, Paul Pu Liang, Zihao Deng, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2022 · 2022
Cited alongside, same era.
Transformerlens
Neel Nanda and Joseph Bloom. 2022 · 2022
Cited alongside, same era.
Generating perturbation-based explanations with robustness to out-of-distribution data
Luyu Qiu, Yi Yang, Caleb Chen Cao, Yueyuan Zheng, Hilary Hei Ting Ngai, Janet Hui-wen Hsiao, and Lei Chen. 2022 · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022 · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Cited alongside, same era.
Scaling and evaluating sparse autoencoders
Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu. 2024 · 2024
Later among the works it cites.
Overthinking the truth: Understanding how language models process false demonstrations
Danny Halawi, Jean-Stanislas Denain, and Jacob Steinhardt. 2024 · 2024
Later among the works it cites.
Finding nemo: Localizing neurons responsible for memorization in diffusion models
Dominik Hintersdorf, Lukas Struppek, Kristian Kersting, Adam Dziedzic, and Franziska Boenisch. 2024 · 2024
Later among the works it cites.
Dissecting misalignment of multimodal large language models via influence function
Lijie Hu, Chenyang Ren, Huanyi Xie, Khouloud Saadi, Shu Yang, Jingfeng Zhang, and Di Wang. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neuron-level interpretation of deep nlp models: A survey
Hassan Sajjad, Nadir Durrani, and Fahim Dalvi. 2022 · 2022
Cited alongside, same era.
FACTOID: A new dataset for identifying misinformation spreaders and political bias
Flora Sakketou, Joan Plepi, Riccardo Cervero, Henri Jacques Geiss, Paolo Rosso, and Lucie Flek. 2022 · 2022
Cited alongside, same era.
Are vision-language transformers learning multimodal representations? a probing perspective
Emmanuelle Salin, Badreddine Farah, Stéphane Ayache, and Benoit Favre. 2022 · 2022
Cited alongside, same era.
Confident adaptive language modeling
Tal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani, Dara Bahri, Vinh Q. Tran, Yi Tay, and Donald Metzler. 2022 · 2022
Cited alongside, same era.
Taking features out of superposition with sparse autoencoders
Lee Sharkey, Dan Braun, and Beren Millidge. 2022 · 2022
Cited alongside, same era.
What the daam: Interpreting stable diffusion using cross attention
Raphael Tang, Linqing Liu, Akshat Pandey, Zhiying Jiang, Gefei Yang, Karun Kumar, Pontus Stenetorp, Jimmy Lin, and Ferhan Ture. 2022 · 2022
Cited alongside, same era.
Localizing and editing knowledge in text-to-image generative models
Samyadeep Basu, Nanxuan Zhao, Vlad I Morariu, Soheil Feizi, and Varun Manjunatha. 2023 · 2023
Cited alongside, same era.
Mmneuron: Discovering neuron-level domain-specific interpretation in multimodal large language model
Jiahao Huo, Yibo Yan, Boren Hu, Yutao Yue, and Xuming Hu. 2024 · 2024
Later among the works it cites.
H-space sparse autoencoders
Ayodeji Ijishakin, Ming Liang Ang, Levente Baljer, Daniel Chee Hian Tan, Hugo Laurence Fry, Ahmed Abdulaal, Aengus Lynch, and James H. Cole. 2024 · 2024
Later among the works it cites.
Clap4clip: Continual learning with probabilistic finetuning for vision-language models
Saurav Jha, Dong Gong, and Lina Yao. 2024 · 2024
Later among the works it cites.
Generating images with multimodal language models
Jing Yu Koh, Daniel Fried, and Russ R Salakhutdinov. 2024 · 2024
Later among the works it cites.
Datainf: Efficiently estimating data influence in lora-tuned llms and diffusion models
Yongchan Kwon, Eric Wu, Kevin Wu, and James Zou. 2024 · 2024
Later among the works it cites.
Modeling caption diversity in contrastive vision-language pretraining
Samuel Lavoie, Polina Kirichenko, Mark Ibrahim, Mahmoud Assran, Andrew Gordon Wilson, Aaron Courville, and Nicolas Ballas. 2024 · 2024
Later among the works it cites.
Sparse autoencoders reveal selective remapping of visual concepts during adaptation
Hyesu Lim, Jinho Choi, Jaegul Choo, and Steffen Schneider. 2024 · 2024
Later among the works it cites.
Zihao Lin, Mohammad Beigi, Hongxuan Li, Yufan Zhou, Yuxiang Zhang, Qifan Wang, Wenpeng Yin, and Lifu Huang. 2024 · 2024
Later among the works it cites.
Grace Luo, Trevor Darrell, and Amir Bar. 2024 · 2024
Later among the works it cites.
Towards principled evaluations of sparse autoencoders for interpretability and control
Aleksandar Makelov, George Lange, and Neel Nanda. 2024 · 2024
Later among the works it cites.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller. 2024 · 2024
Later among the works it cites.
Compositional chain-of-thought prompting for large multimodal models
Chancharik Mitra, Brandon Huang, Trevor Darrell, and Roei Herzig. 2024 · 2024
Later among the works it cites.
Influence functions for scalable data attribution in diffusion models
Bruno Mlodozeniec, Runa Eschenhagen, Juhan Bae, Alexander Immer, David Krueger, and Richard Turner. 2024 · 2024
Later among the works it cites.
Towards interpreting visual information processing in vision-language models
Clement Neo, Luke Ong, Philip Torr, Mor Geva, David Krueger, and Fazl Barez. 2024 · 2024
Later among the works it cites.
A concept-based explainability framework for large multimodal models
Jayneel Parekh, Pegah Khayatan, Mustafa Shukor, Alasdair Newson, and Matthieu Cord. 2024 · 2024
Later among the works it cites.
Explaining generative diffusion models via visual analysis for interpretable decision-making process
Ji-Hoon Park, Yeong-Joon Ju, and Seong-Whan Lee. 2024 · 2024
Later among the works it cites.
Data adaptive traceback for vision-language foundation models in image classification
Wenshuo Peng, Kaipeng Zhang, Yue Yang, Hao Zhang, and Yu Qiao. 2024 · 2024
Later among the works it cites.
Beyond logit lens: Contextual embeddings for robust hallucination detection and grounding in vlms
Anirudh Phukan, Divyansh, Harshit Kumar Morj, Vaishnavi, Apoorv Saxena, and Koustava Goswami. 2024 · 2024
Later among the works it cites.
Grounded text-to-image synthesis with attention refocusing
Quynh Phung, Songwei Ge, and Jia-Bin Huang. 2024 · 2024
Later among the works it cites.
Safe-clip: Removing nsfw concepts from vision-and-language models
Samuele Poppi, Tobia Poppi, Federico Cocchi, Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara. 2024 · 2024
Later among the works it cites.
What factors affect multi-modal in-context learning? an in-depth exploration
Libo Qin, Qiguang Chen, Hao Fei, Zhi Chen, Min Li, and Wanxiang Che. 2024 · 2024
Later among the works it cites.
How and where does clip process negation?
Vincent Quantmeyer, Pablo Mosteiro, and Albert Gatt. 2024 · 2024
Later among the works it cites.
Q-groundcam: Quantifying grounding in vision language models via gradcam
Navid Rajabi and Jana Kosecka. 2024 · 2024
Later among the works it cites.
Jumping ahead: Improving reconstruction fidelity with jumprelu sparse autoencoders
Senthooran Rajamanoharan, Tom Lieberum, Nicolas Sonnerat, Arthur Conmy, Vikrant Varma, János Kramár, and Neel Nanda. 2024 · 2024
Later among the works it cites.
Discover-then-name: Task-agnostic concept bottlenecks via automated concept discovery
Sukrut Rao, Sweta Mahajan, Moritz Böhle, and Bernt Schiele. 2024 · 2024
Later among the works it cites.
Linguistic binding in diffusion models: Enhancing attribute correspondence through attention map alignment
Royi Rassin, Eran Hirsch, Daniel Glickman, Shauli Ravfogel, Yoav Goldberg, and Gal Chechik. 2024 · 2024
Later among the works it cites.
Find: A function description benchmark for evaluating interpretability methods
Sarah Schwettmann, Tamar Shaham, Joanna Materzynska, Neil Chowdhury, Shuang Li, Jacob Andreas, David Bau, and Antonio Torralba. 2024 · 2024
Later among the works it cites.
Lvlm-intrepret: An interpretability tool for large vision-language models
Gabriela Ben Melech Stan, Raanan Yehezkel Rohekar, Yaniv Gurwicz, Matthew Lyle Olson, Anahita Bhiwandiwalla, Estelle Aflalo, Chenfei Wu, Nan Duan, Shao-Yen Tseng, and Vasudev Lal. 2024 · 2024
Later among the works it cites.
A review of multimodal explainable artificial intelligence: Past, present and future
Shilin Sun, Wenbin An, Feng Tian, Fang Nan, Qidong Liu, Jun Liu, Nazaraf Shah, and Ping Chen. 2024 · 2024
Later among the works it cites.
Unpacking sdxl turbo: Interpreting text-to-image models with sparse autoencoders
Viacheslav Surkov, Chris Wendler, Mikhail Terekhov, Justin Deschenaux, Robert West, and Caglar Gulcehre. 2024 · 2024
Later among the works it cites.
Language-specific neurons: The key to multilingual capabilities in large language models
Tianyi Tang, Wenyang Luo, Haoyang Huang, Dongdong Zhang, Xiaolei Wang, Xin Zhao, Furu Wei, and Ji-Rong Wen. 2024 · 2024
Later among the works it cites.
Probing multimodal large language models for global and local semantic representations
Mingxu Tao, Quzhe Huang, Kun Xu, Liwei Chen, Yansong Feng, and Dongyan Zhao. 2024 · 2024
Later among the works it cites.
Chameleon: Mixed-modal early-fusion foundation models
Chameleon Team. 2024 · 2024
Later among the works it cites.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet. transformer circuits thread
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, et al. 2024 · 2024
Later among the works it cites.
Tracrbench: Generating interpretability testbeds with large language models
Hannes Thurnherr and Jérémy Scheurer. 2024 · 2024
Later among the works it cites.
pyvene: A library for understanding and improving PyTorch models via interventions
Zhengxuan Wu, Atticus Geiger, Aryaman Arora, Jing Huang, Zheng Wang, Noah D. Goodman, Christopher D. Manning, and Christopher Potts. 2024 · 2024
Later among the works it cites.
Data attribution for diffusion models: Timestep-induced bias in influence estimation
Tong Xie, Haoyu Li, Andrew Bai, and Cho-Jui Hsieh. 2024b · 2024
Later among the works it cites.
Zhipei Xu, Xuanyu Zhang, Runyi Li, Zecheng Tang, Qing Huang, and Jian Zhang. 2024 · 2024
Later among the works it cites.
Unraveling the shift of visual information flow in MLLMs: From phased interaction to efficient inference
Hao Yin, Guangzong Si, and Zilei Wang. 2024 · 2024
Later among the works it cites.
Neuron-level knowledge attribution in large language models
Zeping Yu and Sophia Ananiadou. 2024b · 2024
Later among the works it cites.
Understanding and mitigating compositional issues in text-to-image generative models
Arman Zarei, Keivan Rezaei, Samyadeep Basu, Mehrdad Saberi, Mazda Moayeri, Priyatham Kattakinda, and Soheil Feizi. 2024 · 2024
Later among the works it cites.
Visual in-context learning for large vision-language models
Yucheng Zhou, Xiang Li, Qianning Wang, and Jianbing Shen. 2024 · 2024
Later among the works it cites.
Chenyi Zhuang, Ying Hu, and Pan Gao. 2024 · 2024
Later among the works it cites.
Safety fine-tuning at (almost) no cost: A baseline for vision large language models
Yongshuo Zong, Ondrej Bohdal, Tingyang Yu, Yongxin Yang, and Timothy Hospedales. 2024 · 2024
Later among the works it cites.
Multi-concept editing using task arithmetic
Niv Cohen, Nicky Kriplani, Benjamin Feuer, Yuval Lemberg, and Chinmay Hegde. 2025 · 2025
Closest in time.
Saeuron: Interpretable concept unlearning in diffusion models with sparse autoencoders
Bartosz Cywiński and Kamil Deja. 2025 · 2025
Closest in time.
Case study: Interpreting, manipulating, and controlling clip with sparse autoencoders
Gytis Daujotas. 2024 · 2025
Closest in time.
Towards neuron attributions in multi-modal large language models
Junfeng Fang, Zac Bi, Ruipeng Wang, Houcheng Jiang, Yuan Gao, Kun Wang, An Zhang, Jie Shi, Xiang Wang, and Tat-Seng Chua. 2025 · 2025
Closest in time.
Concept sliders: Lora adaptors for precise control in diffusion models
Rohit Gandikota, Joanna Materzyńska, Tingrui Zhou, Antonio Torralba, and David Bau. 2025 · 2025
Closest in time.
Chancharik Mitra, Brandon Huang, Tianning Chai, Zhiqiu Lin, Assaf Arbelle, Rogerio Feris, Leonid Karlinsky, Trevor Darrell, Deva Ramanan, and Roei Herzig. 2025 · 2025
Closest in time.
Feature discovery in audio models: A whisper case study
Konstantine Sadov. 2024 · 2025
Closest in time.
Cross-modal safety mechanism transfer in large vision-language models
Shicheng Xu, Liang Pang, Yunchang Zhu, Huawei Shen, and Xueqi Cheng. 2025 · 2025
Closest in time.