Fetching the paper…
Reading the bibliography…
Traditionally, data selection has been studied in settings where all samples from prospective sources are fully revealed to a machine learning developer.
On the translocation of masses
Leonid V Kantorovich · 1942
Earlier work this paper cites.
On information and sufficiency
Solomon Kullback and Richard A Leibler · 1951
Earlier work this paper cites.
Nonlinear programming
Dimitri P Bertsekas · 1997
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
The mnist database of handwritten digits
Yann LeCun · 1998
Earlier work this paper cites.
Hierarchical clustering via joint between-within distances: Extending ward’s minimum variance method
Gabor J Szekely, Maria L Rizzo, et al · 2005
Earlier work this paper cites.
Optimal transport: old and new
Cédric Villani · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
On the kantorovich–rubinstein theorem
David A Edwards · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Active learning: Synthesis lectures on artificial intelligence and machine learning
Burr Settles · 2012
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
Marco Cuturi · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang · 2017
Earlier work this paper cites.
Practical coreset constructions for machine learning
Olivier Bachem, Mario Lucic, and Andreas Krause · 2017
Cited alongside, same era.
Joint distribution optimal transportation for domain adaptation
Nicolas Courty, Rémi Flamary, Amaury Habrard, and Alain Rakotomamonjy · 2017
Cited alongside, same era.
Learning generative models with sinkhorn divergences
Aude Genevay, Gabriel Peyré, and Marco Cuturi · 2018
Cited alongside, same era.
Wasserstein distance guided representation learning for domain adaptation
Jian Shen, Yanru Qu, Weinan Zhang, and Yong Yu · 2018
Cited alongside, same era.
Improving gans using optimal transport
Tim Salimans, Han Zhang, Alec Radford, and Dimitris Metaxas · 2018
Cited alongside, same era.
Geometric dataset distances via optimal transport
David Alvarez-Melis and Nicolo Fusi · 2020
Later among the works it cites.
Deep learning with noisy labels: Exploring techniques and remedies in medical image analysis
Davood Karimi, Haoran Dou, Simon K Warfield, and Ali Gholipour · 2020
Later among the works it cites.
If you like shapley then you’ll love the core
Tom Yan and Ariel D Procaccia · 2021
Later among the works it cites.
Model performance scaling with multiple data sources
Tatsunori Hashimoto · 2021
Later among the works it cites.
Explaining neural scaling laws
Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma · 2021
Later among the works it cites.
Beta shapley: a unified and noise-reduced data valuation framework for machine learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards efficient data valuation based on the shapley value
Ruoxi Jia, David Dao, Boxin Wang, Frances Ann Hubis, Nick Hynes, Nezihe Merve Gürel, Bo Li, Ce Zhang, Dawn Song, and Costas J Spanos · 2019
Cited alongside, same era.
Data shapley: Equitable valuation of data for machine learning
Amirata Ghorbani and James Zou · 2019
Cited alongside, same era.
Efficient task-specific data valuation for nearest neighbor algorithms
Ruoxi Jia, David Dao, Boxin Wang, Frances Ann Hubis, Nezihe Merve Gurel, Bo Li, Ce Zhang, Costas J Spanos, and Dawn Song · 2019
Cited alongside, same era.
Interpolating between optimal transport and mmd using sinkhorn divergences
Jean Feydy, Thibault Séjourné, François-Xavier Vialard, Shun-ichi Amari, Alain Trouvé, and Gabriel Peyré · 2019
Cited alongside, same era.
Universal lipschitz approximation in bounded depth neural networks
Jeremy EJ Cohen, Todd Huster, and Ra Cohen · 2019
Cited alongside, same era.
Sorting out lipschitz function approximation
Cem Anil, James Lucas, and Roger Grosse · 2019
Cited alongside, same era.
Estimating training data influence by tracing gradient descent
Garima Pruthi, Frederick Liu, Satyen Kale, and Mukund Sundararajan · 2020
Cited alongside, same era.
Yongchan Kwon and James Zou · 2021
Later among the works it cites.
Deepcore: A comprehensive library for coreset selection in deep learning
Chengcheng Guo, Bo Zhao, and Yanbing Bai · 2022
Later among the works it cites.
Beyond neural scaling laws: beating power law scaling via data pruning
Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari S Morcos · 2022
Later among the works it cites.
Davinz: Data valuation using deep neural networks at initialization
Zhaoxuan Wu, Yao Shu, and Bryan Kian Hsiang Low · 2022
Later among the works it cites.
Datamodels: Predicting predictions from training data
Andrew Ilyas, Sung Min Park, Logan Engstrom, Guillaume Leclerc, and Aleksander Madry · 2022
Later among the works it cites.
Agreement-on-the-line: Predicting the performance of neural networks under distribution shift
Christina Baek, Yiding Jiang, Aditi Raghunathan, and J Zico Kolter · 2022
Later among the works it cites.
Revisiting neural scaling laws in language and vision
Ibrahim M Alabdulmohsin, Behnam Neyshabur, and Xiaohua Zhai · 2022
Later among the works it cites.
Dataset security for machine learning: Data poisoning, backdoor attacks, and defenses
Micah Goldblum, Dimitris Tsipras, Chulin Xie, Xinyun Chen, Avi Schwarzschild, Dawn Song, Aleksander Mądry, Bo Li, and Tom Goldstein · 2022
Later among the works it cites.
Efficiently computing local lipschitz constants of neural networks via bound propagation
Zhouxing Shi, Yihan Wang, Huan Zhang, J Zico Kolter, and Cho-Jui Hsieh · 2022
Later among the works it cites.
How much more data do i need? estimating requirements for downstream tasks
Rafid Mahmood, James Lucas, David Acuna, Daiqing Li, Jonah Philion, Jose M Alvarez, Zhiding Yu, Sanja Fidler, and Marc T Law · 2022
Later among the works it cites.
Lava: Data valuation without pre-specified learning algorithms
Hoang Anh Just, Feiyang Kang, Tianhao Wang, Yi Zeng, Myeongseob Ko, Ming Jin, and Ruoxi Jia · 2023
Closest in time.
Trak: Attributing model behavior at scale
Sung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc, and Aleksander Madry · 2023
Closest in time.