Fetching the paper…
Reading the bibliography…
Data Shapley provides a principled approach to data valuation and plays a crucial role in data-centric machine learning (ML) research.
Statistical problems in assessing methods of medical diagnosis, with special reference to x-ray techniques
Yerushalmy, J · 1947
Earlier work this paper cites.
A value for n-person games
Shapley, L. S · 1953
Earlier work this paper cites.
Value theory without efficiency
Dubey, P., Neyman, A., and Weber, R. J · 1981
Earlier work this paper cites.
On the inverse problem for semivalues of cooperative tu games
Dragan, I. C · 2002
Earlier work this paper cites.
Vehicle classification in distributed sensor networks
Duarte, M. F. and Hu, Y. H · 2004
Earlier work this paper cites.
Polynomial calculation of the shapley value based on sampling
Castro, J., Gómez, D., and Tejada, J · 2009
Earlier work this paper cites.
The comparisons of data mining techniques for the predictive accuracy of probability of default of credit card clients
Yeh, I.-C. and Lien, C.-h · 2009
Earlier work this paper cites.
Analysis of boolean functions
O’Donnell, R · 2014
Earlier work this paper cites.
Calibrating probability with undersampling for unbalanced classification
Dal Pozzolo, A., Caelen, O., Johnson, R. A., and Bontempi, G · 2015
Earlier work this paper cites.
A new basis and the shapley value
Yokote, K., Funaki, Y., and Kamijo, Y · 2016
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Koh, P. W. and Liang, P · 2017
Earlier work this paper cites.
Representer point selection for explaining deep neural networks
Yeh, C.-K., Kim, J., Yen, I. E.-H., and Ravikumar, P. K · 2018
Earlier work this paper cites.
Data shapley: Equitable valuation of data for machine learning
Ghorbani, A. and Zou, J · 2019
Earlier work this paper cites.
Exploiting negative samples: A catalyst for cohort discovery in healthcare analytics
Zheng, K., Chua, H.-R., Herschel, M., Jagadish, H., Ooi, B. C., and Yip, J. W. L · 2019
Earlier work this paper cites.
A distributional framework for data valuation
Ghorbani, A., Kim, M., and Zou, J · 2020
Earlier work this paper cites.
Estimating training data influence by tracing gradient descent
Pruthi, G., Liu, F., Kale, S., and Sundararajan, M · 2020
Earlier work this paper cites.
Collaborative machine learning with incentive-aware model rewards
Sim, R. H. L., Zhang, Y., Chan, M. C., and Low, B. K. H · 2020
Cited alongside, same era.
Data analysis with shapley values for automatic subject selection in alzheimer’s disease data sets using interpretable machine learning
Bloch, L., Friedrich, C. M., and Initiative, A. D. N · 2021
Cited alongside, same era.
Simple, attack-agnostic defense against targeted training set attacks using cosine similarity
Hammoudeh, Z. and Lowd, D · 2021
Cited alongside, same era.
Efficient computation and analysis of distributional shapley values
Kwon, Y., Rivas, M. A., and Zou, J · 2021
Cited alongside, same era.
Pervasive label errors in test sets destabilize machine learning benchmarks
Northcutt, C. G., Athalye, A., and Mueller, J · 2021
Cited alongside, same era.
Beta shapley: a unified and noise-reduced data valuation framework for machine learning
Kwon, Y. and Zou, J · 2022
Later among the works it cites.
Measuring the effect of training data on deep learning predictions via randomized experiments
Lin, J., Zhang, A., Lécuyer, M., Li, J., Panda, A., and Sen, S · 2022
Later among the works it cites.
Sampling permutations for shapley value estimation
Mitchell, R., Cooper, J., Frank, E., and Holmes, G · 2022
Later among the works it cites.
Data valuation without training of a model
Nohyun, K., Choi, H., and Chung, H. W · 2022
Later among the works it cites.
Understanding influence functions and datamodels via harmonic analysis
Saunshi, N., Gupta, A., Braverman, M., and Arora, S · 2022
Later among the works it cites.
Data valuation in machine learning:“ingredients”, strategies, and open challenges
Sim, R. H. L., Xu, X., and Low, B. K. H · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pandl, K. D., Feiland, F., Thiebes, S., and Sunyaev, A · 2021
Cited alongside, same era.
Deep learning on a data diet: Finding important examples early in training
Paul, M., Ganguli, S., and Dziugaite, G. K · 2021
Cited alongside, same era.
Revisiting methods for finding influential examples
Søgaard, A. et al · 2021
Cited alongside, same era.
Representer point selection via local jacobian expansion for post-hoc classifier explanation of deep neural networks and ensemble models
Sui, Y., Wu, G., and Sanner, S · 2021
Cited alongside, same era.
Data valuation for medical imaging using shapley value and application to a large-scale chest x-ray dataset
Tang, S., Ghorbani, A., Yamashita, R., Rehman, S., Dunnmon, J. A., Zou, J., and Rubin, D. L · 2021
Cited alongside, same era.
Validation free and replication robust volume-based data valuation
Xu, X., Wu, Z., Foo, C. S., and Low, B. K. H · 2021
Cited alongside, same era.
Fundamentals of task-agnostic data valuation
Amiri, M. M., Berdoz, F., and Raskar, R · 2022
Cited alongside, same era.
Later among the works it cites.
Incentivizing collaboration in machine learning via synthetic data rewards
Tay, S. S., Xu, X., Foo, C. S., and Low, B. K. H · 2022
Later among the works it cites.
Davinz: Data valuation using deep neural networks at initialization
Wu, Z., Shu, Y., and Low, B. K. H · 2022
Later among the works it cites.
First is better than last for training data influence
Yeh, C.-K., Taly, A., Sundararajan, M., Liu, F., and Ravikumar, P · 2022
Later among the works it cites.
Simfluence: Modeling the influence of individual training examples by simulating training runs
Guu, K., Webson, A., Pavlick, E., Dixon, L., Tenney, I., and Bolukbasi, T · 2023
Later among the works it cites.
Opendataval: a unified benchmark for data valuation
Jiang, K. F., Liang, W., Zou, J., and Kwon, Y · 2023
Later among the works it cites.
Data-oob: Out-of-bag estimate as a simple and efficient data value
Kwon, Y. and Zou, J · 2023
Later among the works it cites.
Threshold knn-shapley: A linear-time and privacy-friendly approach to data valuation
Wang, J. T., Zhu, Y., Wang, Y.-X., Jia, R., and Mittal, P · 2023
Later among the works it cites.
Impossibility theorems for feature attribution
Bilodeau, B., Jaques, N., Koh, P. W., and Kim, B · 2024
Closest in time.
Efficient data shapley for weighted nearest neighbor algorithms
Wang, J. T., Mittal, P., and Jia, R · 2024
Closest in time.