Fetching the paper…
Reading the bibliography…
The rapid development of Artificial Intelligence (AI) has revolutionized numerous fields, with large language models (LLMs) and computer vision (CV) systems driving advancements in natural language understanding and visual processing, respectively.
A. Dravid, Y. Gandelsman, A. A. Efros, and A. Shocher, “Rosetta neurons: Mining the common units in a model zoo,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 1934–1943
1943
Earlier work this paper cites.
P. Chen, Q. Li, S. Biaz, T. Bui, and A. Nguyen, “gscorecam: What objects is clip looking at?” in Proceedings of the Asian Conference on Computer Vision , 2022, pp. 1959–1975
1975
Earlier work this paper cites.
S. Guillaume, “Designing fuzzy inference systems from data: An interpretability-oriented review,” IEEE Transactions on fuzzy systems , vol. 9, no. 3, pp. 426–443, 2001
2001
Earlier work this paper cites.
S.-M. Zhou and J. Q. Gan, “Low-level interpretability and high-level interpretability: a unified view of data-driven interpretable fuzzy system modelling,” Fuzzy sets and systems , vol. 159, no. 23, pp. 3091–3131, 2008
2008
Earlier work this paper cites.
M. R. Cohen and A. Kohn, “Measuring and interpreting neuronal correlations,” Nature neuroscience , vol. 14, no. 7, pp. 811–819, 2011
2011
Earlier work this paper cites.
C. Szegedy, “Intriguing properties of neural networks,” arXiv preprint arXiv:1312.6199 , 2013
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
M. D. Zeiler and R. Fergus, “Visualizing and understanding convolutional networks,” in Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part I 13 . Springer, 2014, pp. 818–833
2014
Earlier work this paper cites.
A. Mahendran and A. Vedaldi, “Understanding deep image representations by inverting them,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 5188–5196
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
S. Bach, A. Binder, G. Montavon, F. Klauschen, K.-R. Müller, and W. Samek, “On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation,” PloS one , vol. 10, no. 7, p. e0130140, 2015
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Bibal and B. Frénay, “Interpretability of machine learning models and representations: an introduction,” in 24th european symposium on artificial neural networks, computational intelligence and machine learning . CIACO, 2016, pp. 77–82
2016
Earlier work this paper cites.
X. Chen, Y. Duan, R. Houthooft, J. Schulman, I. Sutskever, and P. Abbeel, “Infogan: Interpretable representation learning by information maximizing generative adversarial nets,” Advances in neural information processing systems , vol. 29, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
M. T. Ribeiro, S. Singh, and C. Guestrin, “” why should i trust you?” explaining the predictions of any classifier,” in Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining , 2016, pp. 1135–1144
2016
Earlier work this paper cites.
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba, “Learning deep features for discriminative localization,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 2921–2929
2016
Earlier work this paper cites.
A. Binder, G. Montavon, S. Lapuschkin, K.-R. Müller, and W. Samek, “Layer-wise relevance propagation for neural networks with local renormalization layers,” in Artificial Neural Networks and Machine Learning–ICANN 2016: 25th International Conference on Artificial Neural Networks, Barcelona, Spain, September 6-9, 2016, Proceedings, Part II 25 . Springer, 2016, pp. 63–71
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
S. Chakraborty, R. Tomsett, R. Raghavendra, D. Harborne, M. Alzantot, F. Cerutti, M. Srivastava, A. Preece, S. Julier, R. M. Rao et al. , “Interpretability of deep learning models: A survey of results,” in 2017 IEEE smartworld, ubiquitous intelligence & computing, advanced & trusted computed, scalable computing & communications, cloud & big data computing, Internet of people and smart city innovation (smartworld/SCALCOM/UIC/ATC/CBDcom/IOP/SCI) . IEEE, 2017, pp. 1–6
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
M. Sundararajan, A. Taly, and Q. Yan, “Axiomatic attribution for deep networks,” in International conference on machine learning . PMLR, 2017, pp. 3319–3328
2017
Earlier work this paper cites.
D. Bau, B. Zhou, A. Khosla, A. Oliva, and A. Torralba, “Network dissection: Quantifying interpretability of deep visual representations,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 6541–6549
2017
Earlier work this paper cites.
A. Shrikumar, P. Greenside, and A. Kundaje, “Learning important features through propagating activation differences,” in International conference on machine learning . PMlR, 2017, pp. 3145–3153
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 618–626
2017
Earlier work this paper cites.
G. Montavon, S. Lapuschkin, A. Binder, W. Samek, and K.-R. Müller, “Explaining nonlinear classification decisions with deep taylor decomposition,” Pattern recognition , vol. 65, pp. 211–222, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Voulodimos, N. Doulamis, A. Doulamis, and E. Protopapadakis, “Deep learning for computer vision: A brief review,” Computational intelligence and neuroscience , vol. 2018, no. 1, p. 7068349, 2018
2018
Earlier work this paper cites.
A. Adadi and M. Berrada, “Peeking inside the black-box: a survey on explainable artificial intelligence (xai),” IEEE access , vol. 6, pp. 52 138–52 160, 2018
2018
Earlier work this paper cites.
Q. Zhang, Y. N. Wu, and S.-C. Zhu, “Interpretable convolutional neural networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 8827–8836
2018
Earlier work this paper cites.
Z. C. Lipton, “The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery.” Queue , vol. 16, no. 3, pp. 31–57, 2018
2018
Earlier work this paper cites.
F. K. Došilović, M. Brčić, and N. Hlupić, “Explainable artificial intelligence: A survey,” in 2018 41st International convention on information and communication technology, electronics and microelectronics (MIPRO) . IEEE, 2018, pp. 0210–0215
2018
Earlier work this paper cites.
L. H. Gilpin, D. Bau, B. Z. Yuan, A. Bajwa, M. Specter, and L. Kagal, “Explaining explanations: An overview of interpretability of machine learning,” in 2018 IEEE 5th International Conference on data science and advanced analytics (DSAA) . IEEE, 2018, pp. 80–89
2018
Earlier work this paper cites.
Q.-s. Zhang and S.-C. Zhu, “Visual interpretability for deep learning: a survey,” Frontiers of Information Technology & Electronic Engineering , vol. 19, no. 1, pp. 27–39, 2018
2018
Earlier work this paper cites.
D. H. Park, L. A. Hendricks, Z. Akata, A. Rohrbach, B. Schiele, T. Darrell, and M. Rohrbach, “Multimodal explanations: Justifying decisions and pointing to the evidence,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 8779–8788
2018
Earlier work this paper cites.
A. B. Zadeh, P. P. Liang, S. Poria, E. Cambria, and L.-P. Morency, “Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2018, pp. 2236–2246
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
B. Zhou, Y. Sun, D. Bau, and A. Torralba, “Interpretable basis decomposition for visual explanation,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 119–134
2018
Earlier work this paper cites.
B. Kim, M. Wattenberg, J. Gilmer, C. Cai, J. Wexler, F. Viegas et al. , “Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav),” in International conference on machine learning . PMLR, 2018, pp. 2668–2677
2018
Earlier work this paper cites.
B. Zhou, D. Bau, A. Oliva, and A. Torralba, “Interpreting deep visual representations via network dissection,” IEEE transactions on pattern analysis and machine intelligence , vol. 41, no. 9, pp. 2131–2145, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
C. Olah, A. Satyanarayan, I. Johnson, S. Carter, L. Schubert, K. Ye, and A. Mordvintsev, “The building blocks of interpretability,” Distill , vol. 3, no. 3, p. e10, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
X. Liu, X. Wang, and S. Matwin, “Improving the interpretability of deep neural networks with knowledge distillation,” in 2018 IEEE International Conference on Data Mining Workshops (ICDMW) . IEEE, 2018, pp. 905–912
2018
Earlier work this paper cites.
L. Zhen, P. Hu, X. Wang, and D. Peng, “Deep supervised cross-modal retrieval,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 10 394–10 403
2019
Earlier work this paper cites.
C. E. Zwilling, A. M. Daugherty, C. H. Hillman, A. F. Kramer, N. J. Cohen, and A. K. Barbey, “Enhanced decision-making through multimodal training,” NPJ science of learning , vol. 4, no. 1, p. 11, 2019
2019
Earlier work this paper cites.
M. Du, N. Liu, and X. Hu, “Techniques for interpretable machine learning,” Communications of the ACM , vol. 63, no. 1, pp. 68–77, 2019
2019
Earlier work this paper cites.
S. Hooker, D. Erhan, P.-J. Kindermans, and B. Kim, “A benchmark for interpretability methods in deep neural networks,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
A. Kanehira, K. Takemoto, S. Inayoshi, and T. Harada, “Multimodal explanations by predicting counterfactuality in videos,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 8594–8602
2019
Earlier work this paper cites.
Y. Hu, W. Zhan, L. Sun, and M. Tomizuka, “Multi-modal probabilistic prediction of interactive behavior via an interpretable model,” in 2019 IEEE Intelligent Vehicles Symposium (IV) . IEEE, 2019, pp. 557–563
2019
Earlier work this paper cites.
D.-K. Nguyen and T. Okatani, “Multi-task learning of hierarchical vision-language representation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 10 492–10 501
2019
Earlier work this paper cites.
R. Fong, M. Patrick, and A. Vedaldi, “Understanding deep networks via extremal perturbations and smooth masks,” in Proceedings of the IEEE/CVF international conference on computer vision , 2019, pp. 2950–2958
2019
Earlier work this paper cites.
A. Kanehira and T. Harada, “Learning to explain with complemental examples,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 8603–8611
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
P. Michel, O. Levy, and G. Neubig, “Are sixteen heads really better than one?” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
B. Van Aken, B. Winter, A. Löser, and F. A. Gers, “How does bert answer questions? a layer-wise analysis of transformer representations,” in Proceedings of the 28th ACM international conference on information and knowledge management , 2019, pp. 1823–1832
2019
Earlier work this paper cites.
I. Tenney, “Bert rediscovers the classical nlp pipeline,” arXiv preprint arXiv:1905.05950 , 2019
2019
Earlier work this paper cites.
B. N. Patro, M. Lunayach, S. Patel, and V. P. Namboodiri, “U-cam: Visual explanation using uncertainty based class activation maps,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 7444–7453
2019
Earlier work this paper cites.
Z. Qi, S. Khorram, and F. Li, “Visualizing deep networks by optimizing with integrated gradients.” in CVPR workshops , vol. 2, 2019, pp. 1–4
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
T. Xue, W. Wang, J. Ma, W. Liu, Z. Pan, and M. Han, “Progress and prospects of multimodal fusion methods in physical human–robot interaction: A review,” IEEE Sensors Journal , vol. 20, no. 18, pp. 10 355–10 370, 2020
2020
Earlier work this paper cites.
A. B. Arrieta, N. Díaz-Rodríguez, J. Del Ser, A. Bennetot, S. Tabik, A. Barbado, S. García, S. Gil-López, D. Molina, R. Benjamins et al. , “Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai,” Information fusion , vol. 58, pp. 82–115, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
H. Kaur, H. Nori, S. Jenkins, R. Caruana, H. Wallach, and J. Wortman Vaughan, “Interpreting interpretability: understanding data scientists’ use of interpretability tools for machine learning,” in Proceedings of the 2020 CHI conference on human factors in computing systems , 2020, pp. 1–14
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
A. A. Ismail, M. Gunady, H. Corrada Bravo, and S. Feizi, “Benchmarking deep learning interpretability in time series predictions,” Advances in neural information processing systems , vol. 33, pp. 6441–6452, 2020
2020
Earlier work this paper cites.
P. Linardatos, V. Papastefanopoulos, and S. Kotsiantis, “Explainable ai: A review of machine learning interpretability methods,” Entropy , vol. 23, no. 1, p. 18, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Y. Liu and T. Tuytelaars, “A deep multi-modal explanation model for zero-shot learning,” IEEE Transactions on Image Processing , vol. 29, pp. 4788–4803, 2020
2020
Earlier work this paper cites.
J. Cao, Z. Gan, Y. Cheng, L. Yu, Y.-C. Chen, and J. Liu, “Behind the scene: Revealing the secrets of pre-trained vision-and-language models,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VI 16 . Springer, 2020, pp. 565–580
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Y.-H. H. Tsai, M. Q. Ma, M. Yang, R. Salakhutdinov, and L.-P. Morency, “Multimodal routing: Improving local and global interpretability of multimodal language analysis,” in Proceedings of the Conference on Empirical Methods in Natural Language Processing. Conference on Empirical Methods in Natural Language Processing , vol. 2020. NIH Public Access, 2020, p. 1823
2020
Earlier work this paper cites.
M. A. Jalal, R. Milner, and T. Hain, “Empirical interpretation of speech emotion perception with attention based model for speech emotion recognition,” in Proceedings of Interspeech 2020 . International Speech Communication Association (ISCA), 2020, pp. 4113–4117
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
D. Bau, J.-Y. Zhu, H. Strobelt, A. Lapedriza, B. Zhou, and A. Torralba, “Understanding the role of individual units in a deep neural network,” Proceedings of the National Academy of Sciences , vol. 117, no. 48, pp. 30 071–30 078, 2020
2020
Earlier work this paper cites.
N. Cammarata, G. Goh, S. Carter, L. Schubert, M. Petrov, and C. Olah, “Curve detectors,” Distill , vol. 5, no. 6, pp. e00 024–003, 2020
2020
Earlier work this paper cites.
C. Olah, N. Cammarata, L. Schubert, G. Goh, M. Petrov, and S. Carter, “Zoom in: An introduction to circuits,” Distill , vol. 5, no. 3, pp. e00 024–001, 2020
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of machine learning research , vol. 21, no. 140, pp. 1–67, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
H. Wang, Z. Wang, M. Du, F. Yang, Z. Zhang, S. Ding, P. Mardziel, and X. Hu, “Score-cam: Score-weighted visual explanations for convolutional neural networks,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops , 2020, pp. 24–25
2020
Earlier work this paper cites.
M. Sundararajan, K. Dhamdhere, and A. Agarwal, “The shapley taylor interaction index,” in International conference on machine learning . PMLR, 2020, pp. 9259–9268
2020
Earlier work this paper cites.
E. Härkönen, A. Hertzmann, J. Lehtinen, and S. Paris, “Ganspace: Discovering interpretable gan controls,” Advances in neural information processing systems , vol. 33, pp. 9841–9850, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
P. W. Koh, T. Nguyen, Y. S. Tang, S. Mussmann, E. Pierson, B. Kim, and P. Liang, “Concept bottleneck models,” in International conference on machine learning . PMLR, 2020, pp. 5338–5348
2020
Earlier work this paper cites.
Y. Huang, H. Xue, B. Liu, and Y. Lu, “Unifying multimodal transformer for bi-directional image and text generation,” in Proceedings of the 29th ACM International Conference on Multimedia , 2021, pp. 1138–1147
2021
Earlier work this paper cites.
S. Chun, S. J. Oh, R. S. De Rezende, Y. Kalantidis, and D. Larlus, “Probabilistic embeddings for cross-modal retrieval,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 8415–8424
2021
Earlier work this paper cites.
S. A. Abdu, A. H. Yousef, and A. Salem, “Multimodal video sentiment analysis using deep learning approaches, a survey,” Information Fusion , vol. 76, pp. 204–226, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
M. Dzabraev, M. Kalashnikov, S. Komkov, and A. Petiushko, “Mdmmt: Multidomain multimodal transformer for video retrieval,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 3354–3363
2021
Earlier work this paper cites.
S. Li, P. Zheng, J. Fan, and L. Wang, “Toward proactive human–robot collaborative assembly: A multimodal transfer-learning-enabled action prediction approach,” IEEE Transactions on Industrial Electronics , vol. 69, no. 8, pp. 8579–8588, 2021
2021
Earlier work this paper cites.
H. Chefer, S. Gur, and L. Wolf, “Transformer interpretability beyond attention visualization,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 782–791
2021
Earlier work this paper cites.
Y. Zhang, P. Tiňo, A. Leonardis, and K. Tang, “A survey on neural network interpretability,” IEEE Transactions on Emerging Topics in Computational Intelligence , vol. 5, no. 5, pp. 726–742, 2021
2021
Earlier work this paper cites.
F.-L. Fan, J. Xiong, M. Li, and G. Wang, “On interpretability of artificial neural networks: A survey,” IEEE Transactions on Radiation and Plasma Medical Sciences , vol. 5, no. 6, pp. 741–760, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
G. Goh, N. Cammarata, C. Voss, S. Carter, M. Petrov, L. Schubert, A. Radford, and C. Olah, “Multimodal neurons in artificial neural networks,” Distill , vol. 6, no. 3, p. e30, 2021
2021
Earlier work this paper cites.
E. Rawls, E. Kummerfeld, and A. Zilverstand, “An integrated multimodal model of alcohol use disorder generated by data-driven causal discovery analysis,” Communications biology , vol. 4, no. 1, p. 435, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Y. Rao, W. Zhao, B. Liu, J. Lu, J. Zhou, and C.-J. Hsieh, “Dynamicvit: Efficient vision transformers with dynamic token sparsification,” Advances in neural information processing systems , vol. 34, pp. 13 937–13 949, 2021
2021
Earlier work this paper cites.
X. Zhu, C. Xu, and D. Tao, “Where and what? examining interpretable disentangled representations,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 5861–5870
2021
Cited alongside, same era.
E. Hernandez, S. Schwettmann, D. Bau, T. Bagashvili, A. Torralba, and J. Andreas, “Natural language descriptions of deep visual features,” in International Conference on Learning Representations , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
L. Schubert, C. Voss, N. Cammarata, G. Goh, and C. Olah, “High-low frequency detectors,” Distill , vol. 6, no. 1, pp. e00 024–005, 2021
2021
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
J. Y. Koh, D. Fried, and R. R. Salakhutdinov, “Generating images with multimodal language models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
H. Hu, K. C. Chan, Y.-C. Su, W. Chen, Y. Li, K. Sohn, Y. Zhao, X. Ben, B. Gong, W. Cohen et al. , “Instruct-imagen: Image generation with multi-modal instruction,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 4754–4763
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Chefer, S. Gur, and L. Wolf, “Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 397–406
2021
Cited alongside, same era.
J. D. Janizek, P. Sturmfels, and S.-I. Lee, “Explaining explanations: Axiomatic feature interactions for deep networks,” Journal of Machine Learning Research , vol. 22, no. 104, pp. 1–54, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
J. Ren, D. Zhang, Y. Wang, L. Chen, Z. Zhou, Y. Chen, X. Cheng, X. Wang, M. Zhou, J. Shi et al. , “Towards a unified game-theoretic view of adversarial perturbations and robustness,” Advances in Neural Information Processing Systems , vol. 34, pp. 3797–3810, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
X. Tang, W. Zhang, Y. Yu, K. Turner, T. Derr, M. Wang, and E. Ntoutsi, “Interpretable visual understanding with cognitive attention network,” in Artificial Neural Networks and Machine Learning–ICANN 2021: 30th International Conference on Artificial Neural Networks, Bratislava, Slovakia, September 14–17, 2021, Proceedings, Part I 30 . Springer, 2021, pp. 555–568
2021
Cited alongside, same era.
E. Wong, S. Santurkar, and A. Madry, “Leveraging sparse linear layers for debuggable deep networks,” in International Conference on Machine Learning . PMLR, 2021, pp. 11 205–11 216
2021
Cited alongside, same era.
P. H. Seo, A. Nagrani, A. Arnab, and C. Schmid, “End-to-end generative pretraining for multimodal video captioning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 17 959–17 968
2022
Cited alongside, same era.
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun et al. , “Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 9556–9567
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Chen, L. Xu, H. Zheng, L. Chen, A. Tolba, L. Zhao, K. Yu, and H. Feng, “Evolution and prospects of foundation models: From large language models to large multimodal models.” Computers, Materials & Continua , vol. 80, no. 2, 2024
2024
Closest in time.
2024
Closest in time.
Y. Yan and J. Lee, “Georeasoner: Reasoning on geospatially grounded context for natural language understanding,” in Proceedings of the 33rd ACM International Conference on Information and Knowledge Management , 2024, pp. 4163–4167
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
J. Liu, L. Han, and J. Ji, “Mcan: multimodal causal adversarial networks for dynamic effective connectivity learning from fmri and eeg data,” IEEE Transactions on Medical Imaging , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
N. Wake, A. Kanehira, K. Sasabuchi, J. Takamatsu, and K. Ikeuchi, “Gpt-4v (ision) for robotics: Multimodal task planning from human demonstration,” IEEE Robotics and Automation Letters , 2024
2024
Closest in time.
P. Sermanet, T. Ding, J. Zhao, F. Xia, D. Dwibedi, K. Gopalakrishnan, C. Chan, G. Dulac-Arnold, S. Maddineni, N. J. Joshi et al. , “Robovqa: Multimodal long-horizon reasoning for robotics,” in 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2024, pp. 645–652
2024
Closest in time.
F. Zhao, C. Zhang, and B. Geng, “Deep multimodal data fusion,” ACM Computing Surveys , vol. 56, no. 9, pp. 1–36, 2024
2024
Closest in time.
H. Zhao, H. Chen, F. Yang, N. Liu, H. Deng, H. Cai, S. Wang, D. Yin, and M. Du, “Explainability for large language models: A survey,” ACM Transactions on Intelligent Systems and Technology , vol. 15, no. 2, pp. 1–38, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
J. Rumbelow, “Model agnostic interpretability,” Ph.D. dissertation, The University of St Andrews, 2024
2024
Closest in time.
2024
Closest in time.
S. Leng, H. Zhang, G. Chen, X. Li, S. Lu, C. Miao, and L. Bing, “Mitigating object hallucinations in large vision-language models through visual contrastive decoding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 13 872–13 882
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
G. Verma, M. Choi, K. Sharma, J. Watson-Daniels, S. Oh, and S. Kumar, “Cross-modal projection in multimodal llms doesn’t really project visual attributes to textual space,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , 2024, pp. 657–664
2024
Closest in time.
2024
Closest in time.
T. R. Shaham, S. Schwettmann, F. Wang, A. Rajaram, E. Hernandez, J. Andreas, and A. Torralba, “A multimodal automated interpretability agent,” in Forty-first International Conference on Machine Learning , 2024
2024
Closest in time.
A. Cuadra, J. Breuch, S. Estrada, D. Ihim, I. Hung, D. Askaryar, M. Hassanien, K. L. Fessele, and J. A. Landay, “Digital forms for all: A holistic multimodal large language model agent for health data entry,” Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies , vol. 8, no. 2, pp. 1–39, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Xu, Y. Zhang, E. Xie, Z. Zhao, Y. Guo, K.-Y. K. Wong, Z. Li, and H. Zhao, “Drivegpt4: Interpretable end-to-end autonomous driving via large language model,” IEEE Robotics and Automation Letters , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Q. Huang, X. Dong, P. Zhang, B. Wang, C. He, J. Wang, D. Lin, W. Zhang, and N. Yu, “Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 13 418–13 427
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Wang, Z. Zhang, X. Cheng, R. Huang, L. Liu, Z. Ye, H. Huang, Y. Zhao, T. Jin, P. Gao et al. , “Freebind: Free lunch in unified multimodal space via knowledge fusion,” in Forty-first International Conference on Machine Learning , 2024
2024
Closest in time.
2024
Closest in time.
N. Evirgen, R. Wang, and X. Chen, “From text to pixels: Enhancing user understanding through text-to-image model explanations,” in Proceedings of the 29th International Conference on Intelligent User Interfaces , 2024, pp. 74–87
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
S. Li, F. Xue, K. Liu, D. Guo, and R. Hong, “Multimodal graph causal embedding for multimedia-based recommendation,” IEEE Transactions on Knowledge and Data Engineering , 2024
2024
Closest in time.
V. Swamy, M. Satayeva, J. Frej, T. Bossy, T. Vogels, M. Jaggi, T. Käser, and M.-A. Hartley, “Multimodn—multimodal, multi-task, interpretable modular networks,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
Z. Zhang, Y. Zhong, R. Ming, H. Hu, J. Sun, Z. Ge, Y. Zhu, and X. Jin, “Disttrain: Addressing model and data heterogeneity with disaggregated training for multimodal large language models,” arXiv e-prints , pp. arXiv–2408, 2024
2024
Closest in time.
2024
Closest in time.
T. Yu, Y. Yao, H. Zhang, T. He, Y. Han, G. Cui, J. Hu, Z. Liu, H.-T. Zheng, M. Sun et al. , “Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 13 807–13 816
2024
Closest in time.
R. Mallick, J. Benois-Pineau, and A. Zemmari, “Ifi: Interpreting for improving: A multimodal transformer with an interpretability technique for recognition of risk events,” in International Conference on Multimedia Modeling . Springer, 2024, pp. 117–131
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
T. Fel, V. Boutin, L. Béthune, R. Cadène, M. Moayeri, L. Andéol, M. Chalvidal, and T. Serre, “A holistic approach to unifying automatic concept extraction and concept importance estimation,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
S. Schwettmann, T. Shaham, J. Materzynska, N. Chowdhury, S. Li, J. Andreas, D. Bau, and A. Torralba, “Find: A function description benchmark for evaluating interpretability methods,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
A. S. Sundar, C.-H. H. Yang, D. M. Chan, S. Ghosh, V. Ravichandran, and P. S. Nidadavolu, “Multimodal attention merging for improved speech recognition and audio event classification,” in 2024 IEEE International Conference on Acoustics, Speech, and Signal Processing Workshops (ICASSPW) . IEEE, 2024, pp. 655–659
2024
Closest in time.
V. Lyberatos, S. Kantarelis, E. Dervakos, and G. Stamou, “Perceptual musical features for interpretable audio tagging,” in 2024 IEEE International Conference on Acoustics, Speech, and Signal Processing Workshops (ICASSPW) . IEEE, 2024, pp. 878–882
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
J.-H. Park, Y.-J. Ju, and S.-W. Lee, “Explaining generative diffusion models via visual analysis for interpretable decision-making process,” Expert Systems with Applications , vol. 248, p. 123231, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Y. Chen, P. Cao, Y. Chen, K. Liu, and J. Zhao, “Journey to the center of the knowledge neurons: Discoveries of language-independent knowledge neurons and degenerate knowledge neurons,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 16, 2024, pp. 17 817–17 825
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
T. Tang, W. Luo, H. Huang, D. Zhang, X. Wang, X. Zhao, F. Wei, and J.-R. Wen, “Language-specific neurons: The key to multilingual capabilities in large language models,” in Annual Meeting of the Association for Computational Linguistics , 2024. [Online]. Available: https://api.semanticscholar.org/CorpusID:268032136
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
C. Si, Z. Huang, Y. Jiang, and Z. Liu, “Freeu: Free lunch in diffusion u-net,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 4733–4743
2024
Closest in time.
M. Kowal, R. P. Wildes, and K. G. Derpanis, “Visual concept connectome (vcc): Open world concept discovery and their interlayer connections in deep models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 10 895–10 905
2024
Closest in time.
Q. Ren, J. Gao, W. Shen, and Q. Zhang, “Where we have arrived in proving the emergence of sparse interaction primitives in dnns,” in The Twelfth International Conference on Learning Representations , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
H. Zhou, H. Zhang, H. Deng, D. Liu, W. Shen, S.-H. Chan, and Q. Zhang, “Explaining generalization power of a dnn using interactive concepts,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 15, 2024, pp. 17 105–17 113
2024
Closest in time.
D. Liu, H. Deng, X. Cheng, Q. Ren, K. Wang, and Q. Zhang, “Towards the difficulty for a deep neural network to learn concepts of different complexities,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
H. Deng, N. Zou, M. Du, W. Chen, G. Feng, Z. Yang, Z. Li, and Q. Zhang, “Unifying fourteen post-hoc attribution methods with taylor interactions,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2024
2024
Closest in time.
G. Chen, Y. Li, X. Liu, Z. Li, E. Al Suradi, D. Wei, and K. Zhang, “Llcp: Learning latent causal processes for reasoning-based video question answer,” in ICLR , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
C. Zhou, P. Liu, P. Xu, S. Iyer, J. Sun, Y. Mao, X. Ma, A. Efrat, P. Yu, L. Yu et al. , “Lima: Less is more for alignment,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
D. Mondal, S. Modi, S. Panda, R. Singh, and G. S. Rao, “Kam-cot: Knowledge augmented multimodal chain-of-thoughts reasoning,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 17, 2024, pp. 18 798–18 806
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
X. Zou, Y. Yan, X. Hao, Y. Hu, H. Wen, E. Liu, J. Zhang, Y. Li, T. Li, Y. Zheng et al. , “Deep learning for cross-domain data fusion in urban computing: Taxonomy, advances, and outlook,” Information Fusion , vol. 113, p. 102606, 2025
2025
Closest in time.
Y. Liu, H. Duan, Y. Zhang, B. Li, S. Zhang, W. Zhao, Y. Yuan, J. Wang, C. He, Z. Liu et al. , “Mmbench: Is your multi-modal model an all-around player?” in European Conference on Computer Vision . Springer, 2025, pp. 216–233
2025
Closest in time.
Y. Fan, X. Ma, R. Wu, Y. Du, J. Li, Z. Gao, and Q. Li, “Videoagent: A memory-augmented multimodal agent for video understanding,” in European Conference on Computer Vision . Springer, 2025, pp. 75–92
2025
Closest in time.
M. Nie, R. Peng, C. Wang, X. Cai, J. Han, H. Xu, and L. Zhang, “Reason2drive: Towards interpretable and chain-based reasoning for autonomous driving,” in European Conference on Computer Vision . Springer, 2025, pp. 292–308
2025
Closest in time.