Fetching the paper…
Reading the bibliography…
As larger deep learning models are hard to interpret, there has been a recent focus on generating explanations of these black-box models.
A theory of the learnable
L. G. Valiant · 1984
Earlier work this paper cites.
Probability in Banach Spaces: isoperimetry and processes , volume 23
M. Ledoux and M. Talagrand · 1991
Earlier work this paper cites.
Stochastic programming , volume 6
P. Kall, S. W. Wallace, and P. Kall · 1994
Earlier work this paper cites.
A pac-style model for learning from labeled and unlabeled data
M.-F. Balcan and A. Blum · 2005
Earlier work this paper cites.
Semi-supervised learning (chapelle, o. et al., eds.; 2006)[book reviews]
O. Chapelle, B. Scholkopf, and A. Zien · 2009
Earlier work this paper cites.
A discriminative model for semi-supervised learning
M.-F. Balcan and A. Blum · 2010
Earlier work this paper cites.
Posterior regularization for structured latent variable models
K. Ganchev, J. Graça, J. Gillenwater, and B. Taskar · 2010
Earlier work this paper cites.
Introduction to stochastic programming
J. R. Birge and F. Louveaux · 2011
Earlier work this paper cites.
Learning with noisy labels
N. Natarajan, I. S. Dhillon, P. K. Ravikumar, and A. Tewari · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Earlier work this paper cites.
Deep learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Earlier work this paper cites.
Harnessing deep neural networks with logic rules
Z. Hu, X. Ma, Z. Liu, E. Hovy, and E. Xing · 2016
Earlier work this paper cites.
Data programming: Creating large training sets, quickly
A. J. Ratner, C. M. De Sa, S. Wu, D. Selsam, and C. Ré · 2016
Earlier work this paper cites.
” why should i trust you?” explaining the predictions of any classifier
M. T. Ribeiro, S. Singh, and C. Guestrin · 2016
Earlier work this paper cites.
Sobolev training for neural networks
W. M. Czarnecki, S. Osindero, M. Jaderberg, G. Swirszcz, and R. Pascanu · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
P. W. Koh and P. Liang · 2017
Cited alongside, same era.
A unified approach to interpreting model predictions
S. M. Lundberg and S.-I. Lee · 2017
Cited alongside, same era.
Snorkel: Rapid training data creation with weak supervision
A. Ratner, S. H. Bach, H. Ehrenberg, J. Fries, S. Wu, and C. Ré · 2017
Cited alongside, same era.
Right for the right reasons: training differentiable models by constraining their explanations
A. S. Ross, M. C. Hughes, and F. Doshi-Velez · 2017
Cited alongside, same era.
Smoothgrad: removing noise by adding noise
D. Smilkov, N. Thorat, B. Kim, F. Viégas, and M. Wattenberg · 2017
Cited alongside, same era.
Axiomatic attribution for deep networks
M. Sundararajan, A. Taly, and Q. Yan · 2017
Cited alongside, same era.
Interpretations are useful: penalizing explanations to align neural networks with prior knowledge
L. Rieger, C. Singh, W. Murdoch, and B. Yu · 2020
Later among the works it cites.
Theoretical analysis of self-training with deep networks on unlabeled data
C. Wei, K. Shen, Y. Chen, and T. Ma · 2020
Later among the works it cites.
Self-training with noisy student improves imagenet classification
Q. Xie, M.-T. Luong, E. Hovy, and Q. V. Le · 2020
Later among the works it cites.
On completeness-aware concept-based explanations in deep neural networks
C.-K. Yeh, B. Kim, S. Arik, C.-L. Li, T. Pfister, and P. Ravikumar · 2020
Later among the works it cites.
Lagrangian duality for constrained deep learning
F. Fioretto, P. Van Hentenryck, T. W. Mak, C. Tran, F. Baldo, and M. Lombardi · 2021
Later among the works it cites.
Improving deep learning interpretability by saliency guided training
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Counterfactual explanations without opening the black box: Automated decisions and the gdpr
S. Wachter, B. Mittelstadt, and C. Russell · 2017
Cited alongside, same era.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
B. Kim, M. Wattenberg, J. Gilmer, C. Cai, J. Wexler, F. Viegas, et al · 2018
Cited alongside, same era.
Beyond word importance: Contextual decomposition to extract interactions from lstms
W. J. Murdoch, P. J. Liu, and B. Yu · 2018
Cited alongside, same era.
Representer point selection for explaining deep neural networks
C.-K. Yeh, J. Kim, I. E.-H. Yen, and P. K. Ravikumar · 2018
Cited alongside, same era.
Counterfactual visual explanations
Y. Goyal, Z. Wu, J. Ernst, D. Batra, D. Parikh, and S. Lee · 2019
Cited alongside, same era.
Functional regularization for representation learning: A unified theoretical perspective
S. Garg and Y. Liang · 2020
Cited alongside, same era.
A. A. Ismail, H. Corrada Bravo, and S. Feizi · 2021
Later among the works it cites.
Wrench: A comprehensive benchmark for weak supervision
J. Zhang, Y. Yu, Y. Li, Y. Wang, Y. Yang, M. Yang, and A. Ratner · 2021
Later among the works it cites.
Impossibility theorems for feature attribution
B. Bilodeau, N. Jaques, P. W. Koh, and B. Kim · 2022
Later among the works it cites.
Self-training converts weak learners to strong learners in mixture models
S. Frei, D. Zou, Z. Chen, and Q. Gu · 2022
Later among the works it cites.
Lecture notes from machine learning theory, 2022
T. Ma · 2022
Later among the works it cites.
Label propagation with weak supervision
R. Pukdee, D. Sam, P. K. Ravikumar, and N. Balcan · 2022
Later among the works it cites.
A theory of PAC learnability under transformation invariances
H. Shao, O. Montasser, and A. Blum · 2022
Later among the works it cites.
Supervising model attention with human explanations for robust natural language inference
J. Stacey, Y. Belinkov, and M. Rei · 2022
Later among the works it cites.
Robust explanation constraints for neural networks
M. R. Wicker, J. Heo, L. Costabello, and A. Weller · 2022
Later among the works it cites.
Concept embedding models
M. E. Zarlenga, P. Barbiero, G. Ciravegna, G. Marra, F. Giannini, M. Diligenti, F. Precioso, S. Melacci, A. Weller, P. Lio, et al · 2022
Later among the works it cites.