Fetching the paper…
Reading the bibliography…
The promise of multimodal models for real-world applications has inspired research in visualizing and understanding their internal mechanics with the end goal of empowering stakeholders to visualize model behavior, perform model debugging, and promote trust in machine learning models.
LXMERT: learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal · 1908
Earlier work this paper cites.
Heterogeneous uncertainty sampling for supervised learning
David D Lewis and Jason Catlett · 1994
Earlier work this paper cites.
A sequential algorithm for training text classifiers
David D Lewis and William A Gale · 1994
Earlier work this paper cites.
Audio-visual speech modeling for continuous speech recognition
Stéphane Dupont and Juergen Luettin · 2000
Earlier work this paper cites.
Large-scale concept ontology for multimedia
Milind Naphade, John R Smith, Jelena Tesic, Shih-Fu Chang, Winston Hsu, Lyndon Kennedy, Alexander Hauptmann, and Jon Curtis · 2006
Earlier work this paper cites.
Predictive learning via rule ensembles
Jerome H Friedman and Bogdan E Popescu · 2008
Earlier work this paper cites.
Detecting statistical interactions with additive groves of trees
Daria Sorokina, Rich Caruana, Mirek Riedewald, and Daniel Fink · 2008
Earlier work this paper cites.
Visualizing higher-layer features of a deep network
Dumitru Erhan, Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2009
Earlier work this paper cites.
Active learning literature survey
Burr Settles · 2009
Earlier work this paper cites.
How to explain individual classification decisions
David Baehrens, Timon Schroeter, Stefan Harmeling, Motoaki Kawanabe, Katja Hansen, and Klaus-Robert Müller · 2010
Earlier work this paper cites.
Computing krippendorff’s alpha-reliability
Klaus Krippendorff · 2011
Earlier work this paper cites.
Interpretation and trust: Designing model-driven visualizations for text analysis
Jason Chuang, Daniel Ramage, Christopher Manning, and Jeffrey Heer · 2012
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2013
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Understanding neural networks through deep visualization
Jason Yosinski, Jeff Clune, Thomas Fuchs, and Hod Lipson · 2015
Earlier work this paper cites.
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Earlier work this paper cites.
Towards transparent ai systems: Interpreting visual question answering models
Yash Goyal, Akrit Mohapatra, Devi Parikh, and Dhruv Batra · 2016
Earlier work this paper cites.
Revisiting visual question answering baselines
Allan Jabri, Armand Joulin, and Laurens van der Maaten · 2016
Earlier work this paper cites.
Mimic-iii, a freely accessible critical care database
Alistair EW Johnson, Tom J Pollard, Lu Shen, H Lehman Li-Wei, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark · 2016
Earlier work this paper cites.
" why should i trust you?" explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
VQA: Visual question answering
Aishwarya Agrawal, Jiasen Lu, Stanislaw Antol, Margaret Mitchell, C. Lawrence Zitnick, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
Gated multimodal units for information fusion
John Arevalo, Thamar Solorio, Manuel Montes-y Gómez, and Fabio A González · 2017
Earlier work this paper cites.
Multimodal sentiment analysis with word-level fusion and reinforcement learning
Minghai Chen, Sen Wang, Paul Pu Liang, Tadas Baltrušaitis, Amir Zadeh, and Louis-Philippe Morency · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Earlier work this paper cites.
A review of affective computing: From unimodal analysis to multimodal fusion
Soujanya Poria, Erik Cambria, Rajiv Bajpai, and Amir Hussain · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian J. Goodfellow, Moritz Hardt, and Been Kim · 2018
Earlier work this paper cites.
Blindfold baselines for embodied qa
Ankesh Anand, Eugene Belilovsky, Kyle Kastner, Hugo Larochelle, and Aaron Courville · 2018
Earlier work this paper cites.
Multimodal machine learning: A survey and taxonomy
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency · 2018
Earlier work this paper cites.
Do explanations make vqa models more predictable to a human?
Arjun Chandrasekaran, Viraj Prabhu, Deshraj Yadav, Prithvijit Chattopadhyay, and Devi Parikh · 2018
Earlier work this paper cites.
Embodied question answering
Abhishek Das, Samyak Datta, Georgia Gkioxari, Stefan Lee, Devi Parikh, and Dhruv Batra · 2018
Earlier work this paper cites.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford · 2018
Cited alongside, same era.
Explaining explanations: An overview of interpretability of machine learning
Leilani H Gilpin, David Bau, Ben Z Yuan, Ayesha Bajwa, Michael Specter, and Lalana Kagal · 2018
Cited alongside, same era.
Women also snowboard: Overcoming bias in captioning models
Lisa Anne Hendricks, Kaylee Burns, Kate Saenko, Trevor Darrell, and Anna Rohrbach · 2018
Cited alongside, same era.
The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery
Zachary C Lipton · 2018
Cited alongside, same era.
Efficient low-rank multimodal fusion with modality-specific factors
Zhun Liu, Ying Shen, Varun Bharadhwaj Lakshminarasimhan, Paul Pu Liang, AmirAli Bagher Zadeh, and Louis-Philippe Morency · 2018
Cited alongside, same era.
Towards faithfully interpretable nlp systems: How should we define and evaluate faithfulness?
Alon Jacovi and Yoav Goldberg · 2020
Later among the works it cites.
Multiplicative interactions and where to find them
Siddhant M Jayakumar, Wojciech M Czarnecki, Jacob Menick, Jonathan Schwarz, Jack Rae, Simon Osindero, Yee Whye Teh, Tim Harley, and Razvan Pascanu · 2020
Later among the works it cites.
What does bert with vision look at?
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 2020
Later among the works it cites.
The explanation game: Explaining machine learning models using shapley values
Luke Merrick and Ankur Taly · 2020
Later among the works it cites.
Learning to deceive with attention-based explanations
Danish Pruthi, Mansi Gupta, Bhuwan Dhingra, Graham Neubig, and Zachary C Lipton · 2020
Later among the works it cites.
Interpretation of machine learning models using shapley values: application to compound potency and multi-target activity predictions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multimodal explanations: Justifying decisions and pointing to the evidence
Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata, Anna Rohrbach, Bernt Schiele, Trevor Darrell, and Marcus Rohrbach · 2018
Cited alongside, same era.
Benchmarking deep learning models on large healthcare datasets
Sanjay Purushotham, Chuizheng Meng, Zhengping Che, and Yan Liu · 2018
Cited alongside, same era.
Detecting statistical interactions from neural network weights
Michael Tsang, Dehua Cheng, and Yan Liu · 2018
Cited alongside, same era.
Neural-symbolic vqa: Disentangling reasoning from vision and language understanding
Kexin Yi, Jiajun Wu, Chuang Gan, Antonio Torralba, Pushmeet Kohli, and Josh Tenenbaum · 2018
Cited alongside, same era.
Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph
AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency · 2018
Cited alongside, same era.
Rubi: Reducing unimodal biases for visual question answering
Remi Cadene, Corentin Dancette, Matthieu Cord, Devi Parikh, et al · 2019
Cited alongside, same era.
Multimodal explanations by predicting counterfactuality in videos
Atsushi Kanehira, Kentaro Takemoto, Sho Inayoshi, and Tatsuya Harada · 2019
Cited alongside, same era.
Raquel Rodríguez-Pérez and Jürgen Bajorath · 2020
Later among the works it cites.
Rethinking the role of gradient-based attribution methods for model interpretability
Suraj Srinivas and Francois Fleuret · 2020
Later among the works it cites.
The many shapley values for model explanation
Mukund Sundararajan and Amir Najmi · 2020
Later among the works it cites.
The language interpretability tool: Extensible, interactive visualizations and analysis for nlp models
Ian Tenney, James Wexler, Jasmijn Bastings, Tolga Bolukbasi, Andy Coenen, Sebastian Gehrmann, Ellen Jiang, Mahima Pushkarna, Carey Radebaugh, Emily Reif, et al · 2020
Later among the works it cites.
Nbdt: Neural-backed decision tree
Alvin Wan, Lisa Dunlap, Daniel Ho, Jihan Yin, Scott Lee, Suzanne Petryk, Sarah Adel Bargal, and Joseph E Gonzalez · 2020
Later among the works it cites.
Explain, edit, and understand: Rethinking user study design for evaluating model explanations
Siddhant Arora, Danish Pruthi, Norman Sadeh, William W Cohen, Zachary C Lipton, and Graham Neubig · 2021
Later among the works it cites.
Cooperative learning for multi-view analysis
Daisy Yi Ding and Robert Tibshirani · 2021
Later among the works it cites.
Vision-and-language or vision-for-language? on cross-modal influence in multimodal transformers
Stella Frank, Emanuele Bugliarello, and Desmond Elliott · 2021
Later among the works it cites.
Decoupling the role of data, attention, and losses in multimodal transformers
Lisa Anne Hendricks, John Mellor, Rosalia Schneider, Jean-Baptiste Alayrac, and Aida Nematzadeh · 2021
Later among the works it cites.
Transformer is all you need: Multimodal multitask learning with a unified transformer
Ronghang Hu and Amanpreet Singh · 2021
Later among the works it cites.
Mdetr–modulated detection for end-to-end multi-modal understanding
Aishwarya Kamath, Mannat Singh, Yann LeCun, Ishan Misra, Gabriel Synnaeve, and Nicolas Carion · 2021
Later among the works it cites.
Vilt: Vision-and-language transformer without convolution or region supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim · 2021
Later among the works it cites.
Andreas Madsen, Nicholas Meade, Vaibhav Adlakha, and Siva Reddy · 2021
Later among the works it cites.
Seeing past words: Testing the cross-modal capabilities of pretrained v&l models on counting tasks
Letitia Parcalabescu, Albert Gatt, Anette Frank, and Iacer Calixto · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Do input gradients highlight discriminative features?
Harshay Shah, Prateek Jain, and Praneeth Netrapalli · 2021
Later among the works it cites.
M2lens: Visualizing and explaining multimodal models for sentiment analysis
Xingbo Wang, Jianben He, Zhihua Jin, Muqiao Yang, Yong Wang, and Huamin Qu · 2021
Later among the works it cites.
Leveraging sparse linear layers for debuggable deep networks
Eric Wong, Shibani Santurkar, and Aleksander Madry · 2021
Later among the works it cites.
Vl-interpret: An interactive visualization tool for interpreting vision-language transformers
Estelle Aflalo, Meng Du, Shao-Yen Tseng, Yongfei Liu, Chenfei Wu, Nan Duan, and Vasudev Lal · 2022
Closest in time.
A comparative study of faithfulness metrics for model interpretability methods
Chun Sik Chan, Huanqi Kong, and Liang Guanqing · 2022
Closest in time.
Interpretable machine learning: Moving from mythos to diagnostics
Valerie Chen, Jeffrey Li, Joon Sik Kim, Gregory Plumb, and Ameet Talwalkar · 2022
Closest in time.
Framework for evaluating faithfulness of local explanations
Sanjoy Dasgupta, Nave Frost, and Michal Moshkovitz · 2022
Closest in time.
The disagreement problem in explainable machine learning: A practitioner’s perspective
Satyapriya Krishna, Tessa Han, Alex Gu, Javin Pombra, Shahin Jabbari, Steven Wu, and Himabindu Lakkaraju · 2022
Closest in time.
Image retrieval from contextual descriptions
Benno Krojer, Vaibhav Adlakha, Vibhav Vineet, Yash Goyal, Edoardo Ponti, and Siva Reddy · 2022
Closest in time.
Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency · 2022
Closest in time.
Dime: Fine-grained interpretations of multimodal models via disentangled local explanations
Yiwei Lyu, Paul Pu Liang, Zihao Deng, Ruslan Salakhutdinov, and Louis-Philippe Morency · 2022
Closest in time.
Winoground: Probing vision and language models for visio-linguistic compositionality
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross · 2022
Closest in time.