Fetching the paper…
Reading the bibliography…
Model explanations provide transparency into a trained machine learning model's blackbox behavior to a model builder.
Interpretable and Differentially Private Predictions
Frederik Harder, Matthias Bauer, and Mijung Park. 2020 · 1906
Earlier work this paper cites.
Fooling LIME and SHAP: Adversarial Attacks on Post hoc Explanation Methods
Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju. 2020 · 1911
Earlier work this paper cites.
Model Explanations with Differential Privacy
Neel Patel, Reza Shokri, and Yair Zick. 2020 · 2006
Earlier work this paper cites.
Model extraction from counterfactual explanations
Ulrich Aïvodji, Alexandre Bolot, and Sébastien Gambs. 2020 · 2009
Earlier work this paper cites.
Robust and Stable Black Box Explanations
Himabindu Lakkaraju, Nino Arsov, and Osbert Bastani. 2020 · 2011
Earlier work this paper cites.
Privacy in Pharmacogenetics: An End-to-End Case Study of Personalized Warfarin Dosing. In 23rd USENIX Security Symposium (USENIX Security 14) . USENIX Association, San Diego, CA, 17–32
Matthew Fredrikson, Eric Lantz, Somesh Jha, Simon Lin, David Page, and Thomas Ristenpart. 2014 · 2014
Earlier work this paper cites.
Pufferfish: A Framework for Mathematical Privacy Definitions
Daniel Kifer and Ashwin Machanavajjhala. 2014 · 2014
Earlier work this paper cites.
Model Inversion Attacks That Exploit Confidence Information and Basic Countermeasures. In Proceedings of the 22nd ACM SIGSAC Conference on Computer and Communications Security (Denver, Colorado, USA) (CCS ’15) . Association for Computing Machinery, New York, NY, USA, 1322–1333
Matt Fredrikson, Somesh Jha, and Thomas Ristenpart. 2015 · 2015
Earlier work this paper cites.
Optimized Pre-Processing for Discrimination Prevention. In Advances in Neural Information Processing Systems , I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30. Curran Associates, Inc
Flavio Calmon, Dennis Wei, Bhanukiran Vinzamuri, Karthikeyan Natesan Ramamurthy, and Kush R Varshney. 2017 · 2017
Earlier work this paper cites.
A Unified Approach to Interpreting Model Predictions. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, USA) (NIPS’17) . 4768–4777
Scott M. Lundberg and Su-In Lee. 2017 · 2017
Earlier work this paper cites.
Learning Important Features through Propagating Activation Differences. In Proceedings of the 34th International Conference on Machine Learning - Volume 70 (Sydney, NSW, Australia) (ICML’17) . 3145–3153
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. 2017 · 2017
Earlier work this paper cites.
SmoothGrad: removing noise by adding noise
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda B. Viégas, and Martin Wattenberg. 2017 · 2017
Earlier work this paper cites.
Pufferfish Privacy Mechanisms for Correlated Data. In Proceedings of the 2017 ACM International Conference on Management of Data (Chicago, Illinois, USA) (SIGMOD ’17) . Association for Computing Machinery, New York, NY, USA, 1291–1306
Shuang Song, Yizhen Wang, and Kamalika Chaudhuri. 2017 · 2017
Earlier work this paper cites.
Axiomatic Attribution for Deep Networks. In Proceedings of the 34th International Conference on Machine Learning - Volume 70 (Sydney, NSW, Australia) (ICML’17) . 3319–3328
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Earlier work this paper cites.
Towards better understanding of gradient-based attribution methods for Deep Neural Networks. In International Conference on Learning Representations
Marco Ancona, Enea Ceolini, Cengiz Öztireli, and Markus Gross. 2018 · 2018
Cited alongside, same era.
Property Inference Attacks on Fully Connected Neural Networks Using Permutation Invariant Representations. In Proceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security (Toronto, Canada) (CCS ’18) . Association for Computing Machinery, New York, NY, USA, 619–633
Karan Ganju, Qi Wang, Wei Yang, Carl A. Gunter, and Nikita Borisov. 2018 · 2018
Cited alongside, same era.
AttriGuard: A Practical Defense Against Attribute Inference Attacks via Adversarial Machine Learning. In 27th USENIX Security Symposium (USENIX Security 18) . USENIX Association, Baltimore, MD, 513–529
Jinyuan Jia and Neil Zhenqiang Gong. 2018 · 2018
Cited alongside, same era.
Art. 35 GDPR Data protection impact assessment. In General Data Protection Regulation (GDPR)
European Union Law. 2018 · 2018
Cited alongside, same era.
Does Learning Stable Features Provide Privacy Benefits for Machine Learning Models?. In NeurIPS PPML Workshop
Amit Sharma Divyat Mahajan, Shruti Tople. 2020 · 2020
Later among the works it cites.
Guidance for Regulation of Artificial Intelligence Applications. In Memorandum For The Heads Of Executive Departments And Agencies
White House. 2020 · 2020
Later among the works it cites.
"How Do I Fool You?": Manipulating User Trust via Misleading Black Box Explanations. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (New York, NY, USA) (AIES ’20) . Association for Computing Machinery, New York, NY, USA, 79–85
Himabindu Lakkaraju and Osbert Bastani. 2020 · 2020
Later among the works it cites.
Overlearning Reveals Sensitive Attributes. In International Conference on Learning Representations
Congzheng Song and Vitaly Shmatikov. 2020 · 2020
Later among the works it cites.
Interpretable Deep Learning under Fire. In 29th USENIX Security Symposium (USENIX Security 20) . USENIX Association, 1659–1676
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Privacy Risk in Machine Learning: Analyzing the Connection to Overfitting. In 2018 IEEE 31st Computer Security Foundations Symposium (CSF) . 268–282
Samuel Yeom, Irene Giacomelli, Matt Fredrikson, and Somesh Jha. 2018 · 2018
Cited alongside, same era.
Improved adversarial learning for fair classification
L Elisa Celis and Vijay Keswani. 2019 · 2019
Cited alongside, same era.
Fooling neural network interpretations via adversarial model manipulation
Juyeon Heo, Sunghwan Joo, and Taesup Moon. 2019 · 2019
Cited alongside, same era.
Ethics guidelines for trustworthy AI
High-Level Expert Group on AI. 2019 · 2019
Cited alongside, same era.
Exploiting Unintended Feature Leakage in Collaborative Learning
Luca Melis, Congzheng Song, Emiliano De Cristofaro, and Vitaly Shmatikov. 2019 · 2019
Cited alongside, same era.
Model Reconstruction from Model Explanations. In Proceedings of the Conference on Fairness, Accountability, and Transparency . Association for Computing Machinery, New York, NY, USA, 1–9
Smitha Milli, Ludwig Schmidt, Anca D. Dragan, and Moritz Hardt. 2019 · 2019
Cited alongside, same era.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Cynthia Rudin. 2019 · 2019
Cited alongside, same era.
Interventional Fairness: Causal Database Repair for Algorithmic Fairness. In Proceedings of the 2019 International Conference on Management of Data (Amsterdam, Netherlands) (SIGMOD ’19) . Association for Computing Machinery, New York, NY, USA, 793–810
Babak Salimi, Luke Rodriguez, Bill Howe, and Dan Suciu. 2019 · 2019
Cited alongside, same era.
Xinyang Zhang, Ningfei Wang, Hua Shen, Shouling Ji, Xiapu Luo, and Ting Wang. 2020 · 2020
Later among the works it cites.
DySan: Dynamically Sanitizing Motion Sensor Data Against Sensitive Inferences through Adversarial Networks
Antoine Boutet, Carole Frindel, Sébastien Gambs, Théo Jourdan, and Rosin Claude Ngueveu. 2021 · 2021
Later among the works it cites.
Measuring Data Leakage in Machine-Learning Models with Fisher Information. In Conference on Uncertainty in Artificial Intelligence
Awni Hannun, Chuan Guo, and Laurens van der Maaten. 2021 · 2021
Later among the works it cites.
AI auditing and impact assessment: according to the UK information commissioner’s office
Emre Kazim, Danielle Mendes Thame Denny, and Adriano Koshiyama. 2021 · 2021
Later among the works it cites.
Honest-but-Curious Nets: Sensitive Attributes of Private Inputs Can Be Secretly Coded into the Classifiers’ Outputs. In Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security (CCS ’21)
Mohammad Malekzadeh, Anastasia Borovykh, and Deniz Gündüz. 2021 · 2021
Later among the works it cites.
On the Privacy Risks of Model Explanations. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society (Virtual Event, USA) (AIES ’21) . Association for Computing Machinery, New York, NY, USA, 231–241
Reza Shokri, Martin Strobel, and Yair Zick. 2021 · 2021
Later among the works it cites.
Leakage of Dataset Properties in Multi-Party Machine Learning. In 30th USENIX Security Symposium (USENIX Security 21) . USENIX Association, 2687–2704
Wanrong Zhang, Shruti Tople, and Olga Ohrimenko. 2021 · 2021
Later among the works it cites.
Dikaios: Privacy Auditing of Algorithmic Fairness via Attribute Inference Attacks
Jan Aalmoes, Vasisht Duddu, and Antoine Boutet. 2022 · 2022
Closest in time.
Attribute privacy: Framework and mechanisms. In 2022 ACM Conference on Fairness, Accountability, and Transparency . 757–766
Wanrong Zhang, Olga Ohrimenko, and Rachel Cummings. 2022 · 2022
Closest in time.