Fetching the paper…
Reading the bibliography…
Addressing data integrity challenges, such as unlearning the effects of data poisoning after model training, is necessary for the reliable deployment of machine learning models.
Characterizations of an empirical influence function for detecting influential cases in regression
R Dennis Cook and Sanford Weisberg · 1980
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Towards making systems forget with machine unlearning
Yinzhi Cao and Junfeng Yang · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Certified Defenses for Data Poisoning Attacks
Jacob Steinhardt, Pang Wei W Koh, and Percy S Liang · 2017
Earlier work this paper cites.
Detecting backdoor attacks on deep neural networks by activation clustering
Bryant Chen, Wilka Carvalho, Nathalie Baracaldo, Heiko Ludwig, Benjamin Edwards, Taesung Lee, Ian Molloy, and Biplav Srivastava · 2018
Earlier work this paper cites.
Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks, 2018
Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg · 2018
Earlier work this paper cites.
Spectral Signatures in Backdoor Attacks
Brandon Tran, Jerry Li, and Aleksander Madry · 2018
Earlier work this paper cites.
Deepinspect: A black-box trojan detection and mitigation framework for deep neural networks
Huili Chen, Cheng Fu, Jishen Zhao, and Farinaz Koushanfar · 2019
Earlier work this paper cites.
The potential for artificial intelligence in healthcare
Thomas Davenport and Ravi Kalakota · 2019
Earlier work this paper cites.
Data shapley: Equitable valuation of data for machine learning
Amirata Ghorbani and James Zou · 2019
Earlier work this paper cites.
Making AI forget you: Data deletion in machine learning
Antonio Ginart, Melody Y. Guan, Gregory Valiant, and James Zou · 2019
Earlier work this paper cites.
Badnets: Evaluating backdooring attacks on deep neural networks
Tianyu Gu, Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg · 2019
Earlier work this paper cites.
TABOR: A Highly Accurate Approach to Inspecting and Restoring Trojan Backdoors in AI Systems, 2019
Wenbo Guo, Lun Wang, Xinyu Xing, Min Du, and Dawn Song · 2019
Earlier work this paper cites.
Towards efficient data valuation based on the shapley value
Ruoxi Jia, David Dao, Boxin Wang, Frances Ann Hubis, Nick Hynes, Nezihe Merve Gürel, Bo Li, Ce Zhang, Dawn Song, and Costas J Spanos · 2019
Earlier work this paper cites.
On the accuracy of influence functions for measuring group effects
Pang Wei W Koh, Kai-Siang Ang, Hubert Teo, and Percy S Liang · 2019
Earlier work this paper cites.
Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural Networks
Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li, Bimal Viswanath, Haitao Zheng, and Ben Y. Zhao · 2019
Earlier work this paper cites.
What neural networks memorize and why: Discovering the long tail via influence estimation
Vitaly Feldman and Chiyuan Zhang · 2020
Earlier work this paper cites.
Forgetting outside the box: Scrubbing deep networks of information accessible from input-output observations
Aditya Golatkar, Alessandro Achille, and Stefano Soatto · 2020
Cited alongside, same era.
Certified data removal from machine learning models
Chuan Guo, Tom Goldstein, Awni Hannun, and Laurens van der Maaten · 2020
Cited alongside, same era.
Deep learning in finance and banking: A literature review and classification
Jian Huang, Junyi Chai, and Stella Cho · 2020
Cited alongside, same era.
Estimating training data influence by tracing gradient descent
Garima Pruthi, Frederick Liu, Satyen Kale, and Mukund Sundararajan · 2020
Cited alongside, same era.
Practical Detection of Trojan Neural Networks: Data-Limited and Data-Free Cases
Ren Wang, Gaoyuan Zhang, Sijia Liu, Pin-Yu Chen, Jinjun Xiong, and Meng Wang · 2020
Cited alongside, same era.
Influence functions in deep learning are fragile
Adversarial Unlearning of Backdoors via Implicit Hypergradient
Yi Zeng, Si Chen, Won Park, Z. Morley Mao, Ming Jin, and Ruoxi Jia · 2022
Later among the works it cites.
Can bad teaching induce forgetting? unlearning in deep networks using an incompetent teacher
Vikram S Chundawat, Ayush K Tarun, Murari Mandal, and Mohan Kankanhalli · 2023
Later among the works it cites.
Towards adversarial evaluations for inexact machine unlearning
Shashwat Goel, Ameya Prabhu, Amartya Sanyal, Ser-Nam Lim, Philip Torr, and Ponnurangam Kumaraguru · 2023
Later among the works it cites.
Studying Large Language Model Generalization with Influence Functions
Roger Grosse, Juhan Bae, Cem Anil, Nelson Elhage, Alex Tamkin, Amirhossein Tajdini, Benoit Steiner, Dustin Li, Esin Durmus, Ethan Perez, Evan Hubinger, Kamilė Lukošiūtė, Karina Nguyen, Nicholas Joseph, Sam McCandlish, Jared Kaplan, and Samuel R. Bowman · 2023
Later among the works it cites.
Towards unbounded machine unlearning
Meghdad Kurmanji, Peter Triantafillou, and Eleni Triantafillou · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Samyadeep Basu, Phil Pope, and Soheil Feizi · 2021
Cited alongside, same era.
Machine Unlearning
Lucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia, Adelin Travers, Baiwu Zhang, David Lie, and Nicolas Papernot · 2021
Cited alongside, same era.
Trustworthy AI
Raja Chatila, Virginia Dignum, Michael Fisher, Fosca Giannotti, Katharina Morik, Stuart Russell, and Karen Yeung · 2021
Cited alongside, same era.
Black-Box Detection of Backdoor Attacks With Limited Information and Data
Yinpeng Dong, Xiao Yang, Zhijie Deng, Tianyu Pang, Zihao Xiao, Hang Su, and Jun Zhu · 2021
Cited alongside, same era.
Witches’ brew: Industrial scale data poisoning via gradient matching
Jonas Geiping, Liam H Fowl, W. Ronny Huang, Wojciech Czaja, Gavin Taylor, Michael Moeller, and Tom Goldstein · 2021
Cited alongside, same era.
Adaptive machine unlearning
Varun Gupta, Christopher Jung, Seth Neel, Aaron Roth, Saeed Sharifi-Malvajerdi, and Chris Waites · 2021
Cited alongside, same era.
Descent-to-delete: Gradient-based methods for machine unlearning
Seth Neel, Aaron Roth, and Saeed Sharifi-Malvajerdi · 2021
Cited alongside, same era.
Later among the works it cites.
Trak: Attributing model behavior at scale
Sung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc, and Aleksander Madry · 2023
Later among the works it cites.
Artificial intelligence, machine learning and deep learning in advanced robotics, a review
Mohsen Soori, Behrooz Arezoo, and Roza Dastres · 2023
Later among the works it cites.
Protecting against simultaneous data poisoning attacks
Neel Alex, Shoaib Ahmed Siddiqui, Amartya Sanyal, and David Krueger · 2024
Closest in time.
Training data attribution via approximate unrolled differentation
Juhan Bae, Wu Lin, Jonathan Lorraine, and Roger Grosse · 2024
Closest in time.
Poisoning web-scale training datasets is practical
Nicholas Carlini, Matthew Jagielski, Christopher A Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr · 2024
Closest in time.
Fast machine unlearning without retraining through selective synaptic dampening
Jack Foster, Stefan Schoepf, and Alexandra Brintrup · 2024
Closest in time.
Corrective machine unlearning
Shashwat Goel, Ameya Prabhu, Philip Torr, Ponnurangam Kumaraguru, and Amartya Sanyal · 2024
Closest in time.
Training data influence analysis and estimation: A survey
Zayd Hammoudeh and Daniel Lowd · 2024
Closest in time.
Gex: A flexible method for approximating influence via geometric ensemble
SungYub Kim, Kyungsu Kim, and Eunho Yang · 2024
Closest in time.
Machine Unlearning Fails to Remove Data Poisoning Attacks
Martin Pawelczyk, Jimmy Z. Di, Yiwei Lu, Gautam Kamath, Ayush Sekhari, and Seth Neel · 2024
Closest in time.
Potion: Towards poison unlearning
Stefan Schoepf, Jack Foster, and Alexandra Brintrup · 2024
Closest in time.
If-guide: Influence function-guided detoxification of llms, 2025
Zachary Coalson, Juhan Bae, Nicholas Carlini, and Sanghyun Hong · 2025
Closest in time.
Understanding impact of human feedback via influence functions
Taywon Min, Haeone Lee, Yongchan Kwon, and Kimin Lee · 2025
Closest in time.