Fetching the paper…
Reading the bibliography…
In the rapidly evolving landscape of artificial intelligence, multi-modal large language models are emerging as a significant area of interest.
Ancestral graph markov models
Thomas Richardson and Peter Spirtes · 2002
Earlier work this paper cites.
On the completeness of orientation rules for causal discovery in the presence of latent confounders and selection bias
Jiji Zhang · 2008
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Interpreting cnn knowledge via an explanatory graph
Quanshi Zhang, Ruiming Cao, Feng Shi, Ying Nian Wu, and Song-Chun Zhu · 2018
Earlier work this paper cites.
The what-if tool: Interactive probing of machine learning models
James Wexler, Mahima Pushkarna, Tolga Bolukbasi, Martin Wattenberg, Fernanda Viégas, and Jimbo Wilson · 2019
Earlier work this paper cites.
Towards interpretable object detection by unfolding latent structures
Tianfu Wu and Xi Song · 2019
Earlier work this paper cites.
Interpreting cnns via decision trees
Quanshi Zhang, Yu Yang, Haotian Ma, and Ying Nian Wu · 2019
Earlier work this paper cites.
Iterative causal discovery in the possible presence of latent confounders and selection bias
Raanan Y Rohekar, Shami Nisimov, Yaniv Gurwicz, and Gal Novik · 2021
Earlier work this paper cites.
Covid-transformer: Interpretable covid-19 detection using vision transformer for healthcare
Debaditya Shome, Tejaswini Kar, Sachi Nandan Mohanty, Prayag Tiwari, Khan Muhammad, Abdullah AlTameem, Yazhou Zhang, and Abdul Khader Jilani Saudagar · 2021
Earlier work this paper cites.
Vl-interpret: An interactive visualization tool for interpreting vision-language transformers
Estelle Aflalo, Meng Du, Shao-Yen Tseng, Yongfei Liu, Chenfei Wu, Nan Duan, and Vasudev Lal · 2022
Earlier work this paper cites.
Explaining transformer-based image captioning models: An empirical analysis
Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara · 2022
Earlier work this paper cites.
Towards class interpretable vision transformer with multi-class-tokens
Bowen Dong, Pan Zhou, Shuicheng Yan, and Wangmeng Zuo · 2022
Earlier work this paper cites.
Dime: Fine-grained interpretations of multimodal models via disentangled local explanations
Yiwei Lyu, Paul Pu Liang, Zihao Deng, Ruslan Salakhutdinov, and Louis-Philippe Morency · 2022
Earlier work this paper cites.
xvitcos: Explainable vision transformer based covid-19 screening using radiography
Arnab Kumar Mondal, Arnab Bhattacharjee, Parag Singla, and A. P. Prathosh · 2022
Earlier work this paper cites.
Vision-language transformer for interpretable pathology visual question answering
Usman Naseem, Matloob Khushi, and Jinman Kim · 2022
Earlier work this paper cites.
Clear: Causal explanations from attention in neural recommenders
Shami Nisimov, Raanan Y Rohekar, Yaniv Gurwicz, Guy Koren, and Gal Novik · 2022
Cited alongside, same era.
Focused attention in transformers for interpretable classification of retinal images
Clément Playout, Renaud Duval, Marie Carole Boucher, and Farida Cheriet · 2022
Cited alongside, same era.
Investigation of explainability techniques for multimodal transformers
Krithik Ramesh and Yun Sing Koh · 2022
Cited alongside, same era.
Attention-based interpretability with concept transformers
Mattia Rigotti, Christoph Miksovic, Ioana Giurgiu, Thomas Gschwind, and Paolo Scotton · 2022
Cited alongside, same era.
Explain and improve: Lrp-inference fine-tuning for image captioning models
Jiamei Sun, Sebastian Lapuschkin, Wojciech Samek, and Alexander Binder · 2022
Cited alongside, same era.
Vision diffmask: Faithful interpretation of vision transformers with differentiable patch masking
Angelos Nalmpantis, Apostolos Panagiotopoulos, John Gkountouras, Konstantinos Papakostas, and Wilker Aziz · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Gpt-4v(ision) system card
OpenAi · 2023
Later among the works it cites.
Interpretability-aware vision transformer
Yao Qiang, Chengyin Li, Prashant Khanduri, and Dongxiao Zhu · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mengqi Xue, Qihan Huang, Haofei Zhang, Lechao Cheng, Jie Song, Minghui Wu, and Mingli Song · 2022
Cited alongside, same era.
Swin transformer-based object detection model using explainable meta-learning mining
Ji-Won Baek and Kyungyong Chung · 2023
Cited alongside, same era.
Qwen-vl: A frontier large vision-language model with versatile abilities
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou · 2023
Cited alongside, same era.
A multimodal vision transformer for interpretable fusion of functional and structural neuroimaging data
Yuda Bi, Anees Abrol, Zening Fu, and Vince Calhoun · 2023
Cited alongside, same era.
Explainability in image captioning based on the latent space
Sofiane Elguendouze, Adel Hafiane, Marcilio CP de Souto, and Anaïs Halftermeyer · 2023
Cited alongside, same era.
An interpretable transformer network for the retinal disease classification using optical coherence tomography
Jingzhen He, Junxia Wang, Zeyu Han, Jun Ma, Chongjing Wang, and Meng Qi · 2023
Cited alongside, same era.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung · 2023
Cited alongside, same era.
Evaluation and analysis of hallucination in large vision-language models
Junyang Wang, Yiyang Zhou, Guohai Xu, Pengcheng Shi, Chenlin Zhao, Haiyang Xu, Qinghao Ye, Ming Yan, Ji Zhang, Jihua Zhu, et al · 2023
Later among the works it cites.
https://www.gradio.app/
Gradio · 2024
Closest in time.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2024
Closest in time.
Ifi: Interpreting for improving: A multimodal transformer with an interpretability technique for recognition of risk events
Rupayan Mallick, Jenny Benois-Pineau, and Akka Zemmari · 2024
Closest in time.
Evelyn Mannix and Howard Bondell · 2024
Closest in time.
Causal interpretation of self-attention in pre-trained transformers
Raanan Y Rohekar, Yaniv Gurwicz, and Shami Nisimov · 2024
Closest in time.
Multimodn—multimodal, multi-task, interpretable modular networks
Vinitra Swamy, Malika Satayeva, Jibril Frej, Thierry Bossy, Thijs Vogels, Martin Jaggi, Tanja Käser, and Mary-Anne Hartley · 2024
Closest in time.
Eyes wide shut? exploring the visual shortcomings of multimodal llms
Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, and Saining Xie · 2024
Closest in time.
Hallucination is inevitable: An innate limitation of large language models
Ziwei Xu, Sanjay Jain, and Mohan Kankanhalli · 2024
Closest in time.
Analyzing and mitigating object hallucination in large vision-language models
Yiyang Zhou, Chenhang Cui, Jaehong Yoon, Linjun Zhang, Zhun Deng, Chelsea Finn, Mohit Bansal, and Huaxiu Yao · 2024
Closest in time.