Fetching the paper…
Reading the bibliography…
Data valuation is increasingly used in machine learning (ML) to decide the fair compensation for data owners and identify valuable or harmful data for improving ML models.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
A new measure of rank correlation
Maurice G. Kendall. 1938 · 1938
Earlier work this paper cites.
Gaussian processes for regression
Christopher Williams and Carl Rasmussen. 1995 · 1995
Earlier work this paper cites.
If you like Shapley then you’ll love the core. In Proceedings of the AAAI Conference on Artificial Intelligence . AAAI Press, Palo Alto, California, USA, 5751–5759
Tom Yan and Ariel D Procaccia. 2021 · 1995
Earlier work this paper cites.
Sparse spatial autoregressions
Kelley Pace and Ronald Barry. 1997 · 1997
Earlier work this paper cites.
Algorithms for subset selection in linear regression. In Proceedings of the Fortieth Annual ACM Symposium on Theory of Computing (Victoria, British Columbia, Canada). Association for Computing Machinery, New York, NY, USA, 45–54
Abhimanyu Das and David Kempe. 2008 · 2008
Earlier work this paper cites.
Pearson’s Correlation Coefficient
Wilhelm Kirch (Ed.). 2008 · 2008
Earlier work this paper cites.
Polynomial calculation of the Shapley value based on sampling
Javier Castro, Daniel Gómez, and Juan Tejada. 2009 · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton. 2009 · 2009
Earlier work this paper cites.
Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies . Association for Computational Linguistics, Portland, Oregon, USA, 142–150
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al · 2011
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research [best of the web]
Li Deng. 2012 · 2012
Earlier work this paper cites.
Bounding the estimation error of sampling-based shapley value approximation with/without stratifying
Sasan Maleki, Long Tran-Thanh, Greg Hines, Talal Rahwan, and Alex Rogers. 2013 · 2013
Cited alongside, same era.
Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, Seattle, Washington, USA, 1631–1642
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, and Christopher Potts. 2013 · 2013
Cited alongside, same era.
Computational optimal transport: Complexity by accelerated gradient descent is better than by Sinkhorn’s algorithm. In Proceedings of the 35th International Conference on Machine Learning (ICML) . PMLR, Stockholm, Sweden, 1367–1376
Pavel Dvurechensky, Alexander Gasnikov, and Alexey Kroshnin. 2018 · 2018
Cited alongside, same era.
A new approximation method for the Shapley value applied to the WTC 9/11 terrorist attack
Tjeerd van Campen, Herbert Hamers, Bart Husslage, and Roy Lindelauf. 2018 · 2018
Validation free and replication robust volume-based data valuation. In Advances in Neural Information Processing Systems , Vol. 34. Curran Associates, Inc, Virtual, 10837–10848
Xinyi Xu, Zhaoxuan Wu, Chuan Sheng Foo, and Bryan Kian Hsiang Low. 2021 · 2021
Later among the works it cites.
Beta Shapley: a Unified and Noise-reduced Data Valuation Framework for Machine Learning. In Proceedings of the 25th International Conference on Artificial Intelligence and Statistics (AISTATS) . PMLR, Valencia, Spain, 8780–8802
Yongchan Kwon and James Zou. 2022 · 2022
Later among the works it cites.
Distribution regression with sliced Wasserstein kernels. In Proceedings of the 39th International Conference on Machine Learning (ICML) . PMLR, Maryland, USA, 15501–15523
Dimitri Meunier, Massimiliano Pontil, and Carlo Ciliberto. 2022 · 2022
Later among the works it cites.
Data Valuation in Machine Learning: “Ingredients”, Strategies, and Open Challenges. In Proceedings of the 31st International Joint Conference on Artificial Intelligence (IJCAI) . International Joint Conferences on Artificial Intelligence Organization, Vienna, Austria, 5607–5614
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Data Shapley: Equitable Valuation of Data for Machine Learning. In Proceedings of the 36th International Conference on Machine Learning (ICML) . PMLR, CA, USA, 2242–2251
Amirata Ghorbani and James Zou. 2019 · 2019
Cited alongside, same era.
Towards efficient data valuation based on the Shapley value. In Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics (AISTATS) . PMLR, Okinawa, Japan, 1167–1176
Ruoxi Jia, David Dao, Boxin Wang, Frances Ann Hubis, Nick Hynes, Nezihe Merve Gürel, Bo Li, Ce Zhang, Dawn Song, and Costas J Spanos. 2019 · 2019
Cited alongside, same era.
On the Accuracy of Influence Functions for Measuring Group Effects. In Advances in Neural Information Processing Systems , Vol. 32. Curran Associates, Inc., Vancouver, BC, Canada
Pang Wei W Koh, Kai-Siang Ang, Hubert Teo, and Percy S Liang. 2019 · 2019
Cited alongside, same era.
Geometric dataset distances via optimal transport. In Advances in Neural Information Processing Systems , Vol. 33. Curran Associates, Inc., Virtual, 21428–21439
David Alvarez-Melis and Nicolo Fusi. 2020 · 2020
Cited alongside, same era.
Gaussian processes with multidimensional distribution inputs via optimal transport and Hilbertian embedding
François Bachoc, Alexandra Suvorikova, David Ginsbourger, Jean-Michel Loubes, and Vladimir Spokoiny. 2020 · 2020
Cited alongside, same era.
Collaborative machine learning with incentive-aware model rewards. In Proceedings of the 37th International Conference on Machine Learning (ICML) . PMLR, Vienna, Austria, 8927–8936
Rachael Hwee Ling Sim, Yehong Zhang, Mun Choon Chan, and Bryan Kian Hsiang Low. 2020 · 2020
Cited alongside, same era.
Multidimensional Scaling: Approximation and Complexity. In Proceedings of the 38th International Conference on Machine Learning (ICML) (Proceedings of Machine Learning Research, Vol. 139) . PMLR, Virtual, 2568–2578
Erik Demaine, Adam Hesterberg, Frederic Koehler, Jayson Lynch, and John Urschel. 2021 · 2021
Cited alongside, same era.
Rachael Hwee Ling Sim, Xinyi Xu, and Bryan Kian Hsiang Low. 2022 · 2022
Later among the works it cites.
Improving Cooperative Game Theory-based Data Valuation via Data Utility Learning. In ICLR 2022 Workshop on Socially Responsible Machine Learning . International Conference on Learning Representations, Virtual
Tianhao Wang, Yu Yang, and Ruoxi Jia. 2022 · 2022
Later among the works it cites.
OpenDataVal: a Unified Benchmark for Data Valuation. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track , Vol. 36. Curran Associates Inc., Red Hook, NY, USA, 28624–28647
Kevin Fu Jiang, Weixin Liang, James Zou, and Yongchan Kwon. 2023 · 2023
Later among the works it cites.
LAVA: Data Valuation without Pre-Specified Learning Algorithms. In Proceedings of the 11th International Conference on Learning Representations (ICLR) . OpenReview.net, Kigali, Rwanda
Hoang Anh Just, Feiyang Kang, Tianhao Wang, Yi Zeng, Myeongseob Ko, Ming Jin, and Ruoxi Jia. 2023 · 2023
Later among the works it cites.
Data Banzhaf: A robust data valuation framework for machine learning. In Proceedings of the 26th International Conference on Artificial Intelligence and Statistics (AISTATS) . PMLR, Valencia, Spain, 6388–6421
Jiachen T Wang and Ruoxi Jia. 2023 · 2023
Later among the works it cites.
Approximating the Shapley value without marginal contributions. In Proceedings of the AAAI Conference on Artificial Intelligence . AAAI Press, Vancouver, 13246–13255
Patrick Kolpaczki, Viktor Bengs, Maximilian Muschalik, and Eyke Hüllermeier. 2024 · 2024
Later among the works it cites.
Faster Approximation of Probabilistic and Distributional Values via Least Squares. In Proceedings of the 20th International Conference on Learning Representations (ICLR) . OpenReview.net, Vienna, Austria
Weida Li and Yaoliang Yu. 2024 · 2024
Later among the works it cites.
SAVA: Scalable Learning-Agnostic Data Valuation. In The 13th International Conference on Learning Representations (ICLR) . OpenReview.net, Singapore
Samuel Kessler, Tam Le, and Vu Nguyen. 2025 · 2025
Closest in time.