Fetching the paper…
Reading the bibliography…
A core data-centric learning challenge is the identification of training samples that are detrimental to model performance.
The influence curve and its role in robust estimation
Hampel, F. R · 1974
Earlier work this paper cites.
Residuals and influence in regression
Cook, R. D. and Weisberg, S · 1982
Earlier work this paper cites.
Influence functionals for time series
Martin, R. D. and Yohai, V. J · 1986
Earlier work this paper cites.
Active learning with statistical models
Cohn, D. A., Ghahramani, Z., and Jordan, M. I · 1996
Earlier work this paper cites.
Correlation-based feature selection for machine learning
Hall, M. A · 1999
Earlier work this paper cites.
Distance-based outliers: algorithms and applications
Knorr, E. M., Ng, R. T., and Tucakov, V · 2000
Earlier work this paper cites.
Improving one-class svm for anomaly detection
Li, K.-L., Huang, H.-K., Tian, S.-F., and Xu, W · 2003
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, B. and Brockett, C · 2005
Earlier work this paper cites.
Very sparse random projections
Li, P., Hastie, T. J., and Church, K. W · 2006
Earlier work this paper cites.
Isolation forest
Liu, F. T., Ting, K. M., and Zhou, Z.-H · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C · 2013
Earlier work this paper cites.
Submodularity in data subset selection and active learning
Wei, K., Iyer, R., and Bilmes, J · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Second-order stochastic optimization for machine learning in linear time
Agarwal, N., Bullins, B., and Hazan, E · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Koh, P. W. and Liang, P · 2017
Earlier work this paper cites.
Feature selection in machine learning: A new perspective
Cai, J., Luo, J., Wang, S., and Yang, S · 2018
Earlier work this paper cites.
Efficient task specific data valuation for nearest neighbor algorithms
Jia, R., Dao, D., Wang, B., Hubis, F. A., Gurel, N. M., Li, B., Zhang, C., Spanos, C., and Song, D · 2018
Earlier work this paper cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2018
Earlier work this paper cites.
Representer point selection for explaining deep neural networks
Yeh, C.-K., Kim, J., Yen, I. E.-H., and Ravikumar, P. K · 2018
Earlier work this paper cites.
Input similarity from the neural network perspective
Charpiat, G., Girard, N., Felardos, L., and Tarabalka, Y · 2019
Earlier work this paper cites.
Data shapley: Equitable valuation of data for machine learning
Ghorbani, A. and Zou, J · 2019
Earlier work this paper cites.
Towards efficient data valuation based on the shapley value
Jia, R., Dao, D., Wang, B., Hubis, F. A., Hynes, N., Gürel, N. M., Li, B., Zhang, C., Song, D., and Spanos, C. J · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Cited alongside, same era.
Identifying mislabeled instances in classification datasets
Müller, N. M. and Markert, K · 2019
Cited alongside, same era.
Discriminative jackknife: Quantifying uncertainty in deep learning via higher-order influence functions
Alaa, A. and Van Der Schaar, M · 2020
Cited alongside, same era.
What neural networks memorize and why: Discovering the long tail via influence estimation
Feldman, V. and Zhang, C · 2020
Cited alongside, same era.
Scaling up influence functions
Schioppa, A., Zablotskaia, P., Vilar, D., and Sokolov, A · 2022
Later among the works it cites.
Learning with noisy labels revisited: A study using real-world human annotations
Wei, J., Zhu, Z., Cheng, H., Liu, T., Niu, G., and Liu, Y · 2022
Later among the works it cites.
Dataset pruning: Reducing training data by examining generalization influence
Yang, S., Xie, Z., Peng, H., Xu, M., Sun, M., and Li, P · 2022
Later among the works it cites.
Make every example count: On the stability and utility of self-influence for learning from noisy nlp datasets
Bejan, I., Sokolov, A., and Filippova, K · 2023
Later among the works it cites.
Robust fair clustering: A novel fairness attack and defense framework
Chhabra, A., Li, P., Mohapatra, P., and Liu, H · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Explaining black box predictions and unveiling data artifacts through influence functions
Han, X., Wallace, B. C., and Tsvetkov, Y · 2020
Cited alongside, same era.
Estimating training data influence by tracing gradient descent
Pruthi, G., Liu, F., Kale, S., and Sundararajan, M · 2020
Cited alongside, same era.
Image classification with deep learning in the presence of noisy labels: A survey
Algan, G. and Ulusoy, I · 2021
Cited alongside, same era.
Hydra: Hypergradient data relevance analysis for interpreting deep neural networks
Chen, Y., Li, B., Yu, H., Wu, P., and Miao, C · 2021
Cited alongside, same era.
Retiring adult: New datasets for fair machine learning
Ding, F., Hardt, M., Miller, J., and Schmidt, L · 2021
Cited alongside, same era.
Retrieve: Coreset selection for efficient and robust semi-supervised learning
Killamsetty, K., Zhao, X., Chen, F., and Iyer, R · 2021
Cited alongside, same era.
Resolving training biases via influence-based data relabeling
Kong, S., Shen, Y., and Huang, L · 2021
Cited alongside, same era.
Dai, Z. and Gifford, D. K · 2023
Later among the works it cites.
Studying large language model generalization with influence functions
Grosse, R., Bae, J., Anil, C., Elhage, N., Tamkin, A., Tajdini, A., Steiner, B., Li, D., Durmus, E., Perez, E., et al · 2023
Later among the works it cites.
Efficient Data Subset Selection to Generalize Training Across Models: Transductive and Inductive Networks
Jain, E., Nandy, T., Aggarwal, G., Tendulkar, A. V., Iyer, R. K., and De, A · 2023
Later among the works it cites.
Learning antidote data to individual unfairness
Li, P., Xia, E., and Liu, H · 2023
Later among the works it cites.
Deeper understanding of black-box predictions via generalized influence functions
Lyu, H., Jang, J., Ryu, S., and Yang, H. J · 2023
Later among the works it cites.
Dmlr: Data-centric machine learning research–past, present and future
Oala, L., Maskey, M., Bat-Leah, L., Parrish, A., Gürel, N. M., Kuo, T.-S., Liu, Y., Dror, R., Brajovic, D., Yao, X., et al · 2023
Later among the works it cites.
Trak: Attributing model behavior at scale
Park, S. M., Georgiev, K., Ilyas, A., Leclerc, G., and Madry, A · 2023
Later among the works it cites.
Add-remove-or-relabel: Practitioner-friendly bias mitigation via influential fairness
Richardson, B., Sattigeri, P., Wei, D., Ramamurthy, K. N., Varshney, K., Dhurandhar, A., and Gilbert, J. E · 2023
Later among the works it cites.
Self-influence guided data reweighting for language model pre-training
Thakkar, M., Bolukbasi, T., Ganapathy, S., Vashishth, S., Chandar, S., and Talukdar, P · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Data-centric artificial intelligence: A survey
Zha, D., Bhat, Z. P., Lai, K.-H., Yang, F., Jiang, Z., Zhong, S., and Hu, X · 2023
Later among the works it cites.
Training data attribution via approximate unrolled differentation
Bae, J., Lin, W., Lorraine, J., and Grosse, R · 2024
Closest in time.
What Data Benefits My Classifier? Enhancing Model Performance and Interpretability through Influence-Based Data Selection
Chhabra, A., Li, P., Mohapatra, P., and Liu, H · 2024
Closest in time.
Gex: A flexible method for approximating influence via geometric ensemble
Kim, S., Kim, K., and Yang, E · 2024
Closest in time.
DataInf: Efficiently Estimating Data Influence in LoRA-tuned LLMs and Diffusion Models
Kwon, Y., Wu, E., Wu, K., and Zou, J · 2024
Closest in time.
Theoretical and practical perspectives on what influence functions do
Schioppa, A., Filippova, K., Titov, I., and Zablotskaia, P · 2024
Closest in time.
Data pruning via moving-one-sample-out
Tan, H., Wu, S., Du, F., Chen, Y., Wang, Z., Wang, F., and Qi, X · 2024
Closest in time.
Unraveling Indirect In-Context Learning Using Influence Functions
Askari, H., Gupta, S., Tong, T., Wang, F., Chhabra, A., and Chen, M · 2025
Closest in time.