Fetching the paper…
Reading the bibliography…
While deep learning models often lack interpretability, concept bottleneck models (CBMs) provide inherent explanations via their concept representations.
The mnist database of handwritten digits
LeCun, Y. and Cortes, C · 1998
Earlier work this paper cites.
Active learning literature survey
Settles, B · 2009
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y · 2011
Earlier work this paper cites.
Caltech-ucsd birds
Wah, C., Branson, S., Welinder, P., Perona, P., and Belongie, S · 2011
Earlier work this paper cites.
Power to the people: The role of humans in interactive machine learning
Amershi, S., Cakmak, M., Knox, W. B., and Kulesza, T · 2014
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Earlier work this paper cites.
A baseline for detecting misclassified and out-of-distribution examples in neural networks
Hendrycks, D. and Gimpel, K · 2017
Earlier work this paper cites.
Peeking inside the black-box: A survey on explainable artificial intelligence (XAI)
Adadi, A. and Berrada, M · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV)
Kim, B., Wattenberg, M., Gilmer, J., Cai, C. J., Wexler, J., Viégas, F. B., and Sayres, R · 2018
Earlier work this paper cites.
Neural nearest neighbors networks
Plötz, T. and Roth, S · 2018
Earlier work this paper cites.
Gradient based sample selection for online continual learning
Aljundi, R., Lin, M., Goujaud, B., and Bengio, Y · 2019
Earlier work this paper cites.
Ilastik: interactive machine learning for (bio) image analysis
Berg, S., Kutra, D., Kroeger, T., Straehle, C. N., Kausler, B. X., Haubold, C., Schiegg, M., Ales, J., Beier, T., Rudy, M., et al · 2019
Earlier work this paper cites.
Explanation in artificial intelligence: Insights from the social sciences
Miller, T · 2019
Earlier work this paper cites.
Density estimation in representation space to predict model uncertainty
Ramalho, T. and Miranda, M · 2019
Earlier work this paper cites.
Concept bottleneck models
Koh, P. W., Nguyen, T., Tang, Y. S., Mussmann, S., Pierson, E., Kim, B., and Liang, P · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive NLP tasks
Lewis, P. S. H., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., and Kiela, D · 2020
Earlier work this paper cites.
Interpretations are useful: Penalizing explanations to align neural networks with prior knowledge
Rieger, L., Singh, C., Murdoch, W. J., and Yu, B · 2020
Cited alongside, same era.
Making deep neural networks right for the right scientific reasons by interacting with their explanations
Schramowski, P., Stammer, W., Teso, S., Brugger, A., Herbert, F., Shao, X., Luigs, H., Mahlein, A., and Kersting, K · 2020
Cited alongside, same era.
TabCBM: Concept-based interpretable neural networks for tabular data
Zarlenga, M. E., Shams, Z., Nelson, M. E., Kim, B., and Jamnik, M · 2020
Cited alongside, same era.
Toward a unified framework for debugging concept-based models
Bontempelli, A., Giunchiglia, F., Passerini, A., and Teso, S · 2021
Cited alongside, same era.
A survey of uncertainty in deep neural networks
Gawlikowski, J., Tassi, C. R. N., Ali, M., Lee, J., Humt, M., Feng, J., Kruspe, A. M., Triebel, R., Jung, P., Roscher, R., Shahzad, M., Yang, W., Bamler, R., and Zhu, X. X · 2021
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., and Lowe, R · 2022
Later among the works it cites.
Concept bottleneck model with additional unsupervised concepts
Sawada, Y. and Nakamura, K · 2022
Later among the works it cites.
Learning from uncertain concepts via test time interventions
Sheth, I., Rahman, A. A., Sevyeri, L. R., Havaei, M., and Kahou, S. E · 2022
Later among the works it cites.
Interactive disentanglement: Learning concepts by interacting with their prototype representations
Stammer, W., Memmel, M., Schramowski, P., and Kersting, K · 2022
Later among the works it cites.
Post-hoc concept bottleneck models
Yüksekgönül, M., Wang, M., and Zou, J · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Billion-scale similarity search with gpus
Johnson, J., Douze, M., and Jégou, H · 2021
Cited alongside, same era.
Promises and pitfalls of black-box concept learning models
Mahinpei, A., Clark, J., Lage, I., Doshi-Velez, F., and Pan, W · 2021
Cited alongside, same era.
Do concept bottleneck models learn as intended?
Margeloiu, A., Ashman, M., Bhatt, U., Chen, Y., Jamnik, M., and Weller, A · 2021
Cited alongside, same era.
Right for the right concept: Revising neuro-symbolic concepts by interacting with their explanations
Stammer, W., Schramowski, P., and Kersting, K · 2021
Cited alongside, same era.
Meaningfully debugging model mistakes using conceptual counterfactual explanations
Abid, A., Yüksekgönül, M., and Zou, J · 2022
Cited alongside, same era.
Improving language models by retrieving from trillions of tokens
Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., van den Driessche, G., Lespiau, J., Damoc, B., Clark, A., de Las Casas, D., Guy, A., Menick, J., Ring, R., Hennigan, T., Huang, S., Maggiore, L., Jones, C., Cassirer, A., Brock, A., Paganini, M., Irving, G., Vinyals, O., Osindero, S., Simonyan, K., Rae, J. W., Elsen, E., and Sifre, L · 2022
Cited alongside, same era.
Interactive concept bottleneck models
Chauhan, K., Tiwari, R., Freyberg, J., Shenoy, P., and Dvijotham, K · 2022
Cited alongside, same era.
Concept embedding models: Beyond the accuracy-explainability trade-off
Zarlenga, M. E., Barbiero, P., Ciravegna, G., Marra, G., Giannini, F., Diligenti, M., Shams, Z., Precioso, F., Melacci, S., Weller, A., Lió, P., and Jamnik, M · 2022
Later among the works it cites.
A survey on XAI and natural language explanations
Cambria, E., Malandri, L., Mercorio, F., Mezzanzanica, M., and Nobani, N · 2023
Closest in time.
Human uncertainty in concept-based AI systems
Collins, K. M., Barker, M., Zarlenga, M. E., Raman, N., Bhatt, U., Jamnik, M., Sucholutsky, I., Weller, A., and Dvijotham, K · 2023
Closest in time.
A typology for exploring the mitigation of shortcut behaviour
Friedrich, F., Stammer, W., Schramowski, P., and Kersting, K · 2023
Closest in time.
Probabilistic concept bottleneck models
Kim, E., Jung, D., Park, S., Kim, S., and Yoon, S · 2023
Closest in time.
Neuro-symbolic reasoning shortcuts: Mitigation strategies and their limitations
Marconato, E., Teso, S., and Passerini, A · 2023
Closest in time.
Label-free concept bottleneck models
Oikarinen, T. P., Das, S., Nguyen, L. M., and Weng, T · 2023
Closest in time.
Explainable AI (XAI): A systematic meta-survey of current challenges and future opportunities
Saeed, W. and Omlin, C. W · 2023
Closest in time.
A closer look at the intervention procedure of concept bottleneck models
Shin, S., Jo, Y., Ahn, S., and Lee, N · 2023
Closest in time.
Leveraging explanations in interactive machine learning: An overview
Teso, S., Alkan, Ö., Stammer, W., and Daly, E · 2023
Closest in time.