Fetching the paper…
Reading the bibliography…
There has been considerable recent interest in interpretable concept-based models such as Concept Bottleneck Models (CBMs), which first predict human-interpretable concepts and then map them to output classes.
Regularization and variable selection via the elastic net
Hui Zou and Trevor Hastie · 2005
Earlier work this paper cites.
Learning to detect unseen object classes by between-class attribute transfer
Christoph H Lampert, Hannes Nickisch, and Stefan Harmeling · 2009
Earlier work this paper cites.
Attribute and simile classifiers for face verification
Neeraj Kumar, Alexander C Berg, Peter N Belhumeur, and Shree K Nayar · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset, 2011
Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie · 2011
Earlier work this paper cites.
K-sparse autoencoders
Alireza Makhzani and Brendan Frey · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Places: A 10 million image database for scene recognition
Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba · 2017
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
David Bau, Bolei Zhou, Aditya Khosla, Aude Oliva, and Antonio Torralba · 2017
Earlier work this paper cites.
Feature visualization
Chris Olah, Alexander Mordvintsev, and Ludwig Schubert · 2017
Earlier work this paper cites.
Dictionary learning algorithms and applications
Bogdan Dumitrescu and Paul Irofti · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al · 2018
Earlier work this paper cites.
Interpretable basis decomposition for visual explanation
Bolei Zhou, Yiyou Sun, David Bau, and Antonio Torralba · 2018
Earlier work this paper cites.
Towards automatic concept-based explanations
Amirata Ghorbani, James Wexler, James Y Zou, and Been Kim · 2019
Earlier work this paper cites.
Concept Bottleneck Models
Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pierson, Been Kim, and Percy Liang · 2020
Earlier work this paper cites.
Modifying memories in transformer models
Chen Zhu, Ankit Singh Rawat, Manzil Zaheer, Srinadh Bhojanapalli, Daliang Li, Felix Yu, and Sanjiv Kumar · 2020
Cited alongside, same era.
Rewriting a deep generative model
David Bau, Steven Liu, Tongzhou Wang, Jun-Yan Zhu, and Antonio Torralba · 2020
Cited alongside, same era.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Cited alongside, same era.
Invertible concept-based explanations for cnn models with non-negative concept activation vectors
Ruihan Zhang, Prashan Madumal, Tim Miller, Krista A Ehinger, and Benjamin IP Rubinstein · 2021
Cited alongside, same era.
Leveraging sparse linear layers for debuggable deep networks
Eric Wong, Shibani Santurkar, and Aleksander Madry · 2021
Cited alongside, same era.
Label-free Concept Bottleneck Models
Tuomas Oikarinen, Subhro Das, Lam M. Nguyen, and Tsui-Wei Weng · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency, 2023
Sarah Chen James Campbell Phillip Guo Richard Ren Alexander Pan Xuwang Yin Mantas Mazeika Ann-Kathrin Dombrowski Shashwat Goel Nathaniel Li Michael J. Byun Zifan Wang Alex Mallen Steven Basart Sanmi Koyejo Dawn Song Matt Fredrikson Zico Kolter Dan Hendrycks Andy Zou, Long Phan · 2023
Later among the works it cites.
Multi-dimensional concept discovery (MCD): A unifying framework with completeness guarantees
Johanna Vielhaben, Stefan Bluecher, and Nils Strodthoff · 2023
Later among the works it cites.
Overlooked factors in concept-based explanations: Dataset choice, concept learnability, and human capability
Vikram V Ramaswamy, Sunnie SY Kim, Ruth Fong, and Olga Russakovsky · 2023
Later among the works it cites.
Gpt-4 technical report
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andrei Margeloiu, Matthew Ashman, Umang Bhatt, Yanzhi Chen, Mateja Jamnik, and Adrian Weller · 2021
Cited alongside, same era.
Promises and pitfalls of black-box concept learning models
Anita Mahinpei, Justin Clark, Isaac Lage, Finale Doshi-Velez, and Weiwei Pan · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Editing a classifier by rewriting its prediction rules
Shibani Santurkar, Dimitris Tsipras, Mahalaxmi Elango, David Bau, Antonio Torralba, and Aleksander Madry · 2021
Cited alongside, same era.
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, et al · 2022
Cited alongside, same era.
Addressing leakage in concept bottleneck models
Marton Havasi, Sonali Parbhoo, and Finale Doshi-Velez · 2022
Cited alongside, same era.
Glancenets: Interpretable, leak-proof concept-based models
Emanuele Marconato, Andrea Passerini, and Stefano Teso · 2022
Cited alongside, same era.
Language in a bottle: Language model guided concept bottlenecks for interpretable image classification
Yue Yang, Artemis Panagopoulou, Shenghao Zhou, Daniel Jin, Chris Callison-Burch, and Mark Yatskar · 2023
Later among the works it cites.
Concept bottleneck generative models
Aya Abdelsalam Ismail, Julius Adebayo, Hector Corrada Bravo, Stephen Ra, and Kyunghyun Cho · 2023
Later among the works it cites.
TabCBM: Concept-based interpretable neural networks for tabular data
Mateo Espinosa Zarlenga, Zohreh Shams, Michael Edward Nelson, Been Kim, and Mateja Jamnik · 2023
Later among the works it cites.
Visual classification via description from large language models
Sachit Menon and Carl Vondrick · 2023
Later among the works it cites.
Erasing concepts from diffusion models
Rohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, and David Bau · 2023
Later among the works it cites.
Beyond concept bottleneck models: How to make black boxes intervenable?
Ričards Marcinkevičs, Sonia Laguna, Moritz Vandenhirtz, and Julia E Vogt · 2024
Closest in time.
Sparse autoencoders find highly interpretable features in language models
Robert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart, and Lee Sharkey · 2024
Closest in time.
Towards compositionality in concept learning
Adam Stein, Aaditya Naik, Yinjun Wu, Mayur Naik, and Eric Wong · 2024
Closest in time.
Explaining explainability: Understanding concept activation vectors
Angus Nicolson, Lisa Schut, J Alison Noble, and Yarin Gal · 2024
Closest in time.
Do concept bottleneck models obey locality?
Naveen Raman, Mateo Espinosa Zarlenga, Juyeon Heo, and Mateja Jamnik · 2024
Closest in time.
Describing differences in image sets with natural language
Lisa Dunlap, Yuhui Zhang, Xiaohan Wang, Ruiqi Zhong, Trevor Darrell, Jacob Steinhardt, Joseph E Gonzalez, and Serena Yeung-Levy · 2024
Closest in time.